Abstract: In September 2026, GPT-6 Astra reached 99.9% on ARC-AGI-3, a benchmark built to measure skill acquisition in novel environments, and surpassed the median human on action efficiency across 96% of levels. Run through a pre-registered battery of eight stylised facts about human trust behaviour, the same model reproduces none of them. Across eight models from five organisations, no model reproduces more than two of the six facts fixed in advance, and three tie there across a size range spanning 27B open weights to a hosted frontier system. Capability does not order the table, shown within two lineages that hold the laboratory fixed. The failure is not uniform: the parameter that matters most for anyone modelling human trust in AI agents, sharper withdrawal from an automated counterpart after an identical error, is reproduced by no model, and for five of seven the criterion could not have been met at all. The consequence for the ATDP framework is a boundary rather than a verdict: a small, instrument-conditional set of human-human quantities can be calibrated from a proxy validated fact by fact and role by role; the human-AI coupling term cannot, and should remain swept.
On 3 September the ARC Prize Foundation published its evaluation of OpenAI’s GPT-6 Astra on ARC-AGI-3 [1]. The benchmark exists to measure what static evaluations miss. An agent is dropped into an unfamiliar, turn-based environment with no instructions. It has to explore, infer the mechanics, work out what counts as a goal, and plan its way there. Human participants solve 100% of the environments.
The results are worth reading carefully, because the number that gets quoted is a fact about the measurement apparatus as much as about the model.
Reasoning effort Standard harness Provider adapter max 62.7% 98.6% xhigh 59.3% 98.4% high 54.8% 99.9% medium 38.6% 98.4% low 17.5% 98.0% none 35.2% 96.7%
Source: ARC Prize Foundation [1]. The standard harness gives every provider the same minimal interface and leaves the model to decide what to preserve in visible notes. The provider adapter preserves opaque reasoning state between requests.
Three numbers for one model on one evaluation set: 99.9%, 62.7%, 17.5%. The standard-harness column is not even monotone in reasoning effort, with “none” above “low” by seventeen points.
The efficiency result is the more significant one, and it is not an artefact of harness choice. On 96% of the levels it completed, Astra used fewer actions than the median human tested, and about half as many per level on average. The foundation had expected action efficiency to remain a dividing line between humans and machines. It states plainly that it no longer is.
It is also careful about scope, and says so in the post: the environments are deterministic and closed-ended, the format is tightly bounded, and clearing the bar is not a claim about general intelligence.
For the last several months I have been running language models through a different kind of evaluation. Not an abstract environment. A relationship, with a counterpart who lets you down once, and a decision every round about how much to expose yourself again. Astra went through it as soon as I could reach it.
Astra reproduces none of the eight human behaviours it was scored against. It opens at the maximum on every one of thirty runs, with a between-run standard deviation of exactly zero.
That is not a gotcha about a strong model. It is a statement about which behaviours the current evaluation stack measures and which it does not touch, and one of the untouched ones is load-bearing for a question I think matters more than the leaderboard.
I. Why what a model does after a betrayal is a parameter, not a curiosity
Trust between people aggregates into something with measurable consequences. Putnam spent decades showing that the density of reciprocity, generalised trust and civic participation in a community tracks how well its institutions work and how prosperous it becomes, and that the pattern is stable enough to persist across centuries [2, 3]. Coleman located the resource in the structure of relationships rather than in any individual holding it [4]. The trust and growth association has since been documented in economics across several measurement approaches [5, 6]. In earlier work with colleagues at Birkbeck and Messina, we found trust parameters inferred from platform interaction data associated strongly with regional economic indicators [7]. I treat that as motivating rather than settled.
The mechanism has always had humans on both ends. Machines were studied extensively as objects of reliance, and reliance is not the same thing. You rely on a cash machine. You do not trust it in the sense that matters here, because you form no belief about whether it is oriented toward your interests.
What changed is not that the machines acquired intentions. It is that people now form exactly the belief the old framework required. The socio-cognitive model of trust [8] was never a description of the trustee’s inner life. It is a description of the trustor’s belief structure: that the other party has the opportunity to act, the competence to act effectively, and the willingness to act in your interest. Cash machines were excluded not because anyone verified the absence of their intentions, but because people reliably do not form willingness-beliefs about them. They demonstrably do form them about conversational agents.
So agents are inside the mechanism. The open question is what they contribute once they are there, and any dynamical model of that needs the same small set of constants: a rate at which beliefs update on evidence, a magnitude by which a betrayal reduces them, a threshold below which trust is not extended again, and a coupling term saying whether experience with agents feeds the human-human mechanism at all. I take those labels from the ATDP framework [9], which declares the last one unknown and sweeps it.
II. What the dynamics turn on
Of those four constants, three sit on the human-human axis and are in principle recoverable from laboratory work that already exists. The fourth, the coupling term, is not, and it is the one that decides the answer. Here is why.
Social capital accumulates and decays at the same time. Trust successfully extended to someone outside your existing circle adds to the stock. The stock erodes when it is not renewed. A community that has looked stable for decades is one where those two rates happen to balance.
Now let agents take over some share of the interactions that used to do the adding. Not all of them, and not badly. Assume every one of those interactions goes well: the agent is competent, transparent, and does what it said it would. The only question left is whether an interaction with an agent adds as much to the stock as the interaction with a person that it replaced.
If it adds the same, nothing happens and there is nothing to discuss. If it adds less, the balance breaks.
The intuitive expectation is that it breaks gently, and that there is some tolerable level of substitution below which nothing measurable occurs. There is not. A balance is a balance: remove any fraction of the inflow and the stock starts falling from that moment. No threshold has to be crossed first. Decline begins at the first substituted interaction and continues smoothly, and whatever tipping behaviour the system has arrives later and on top of that.
What makes this hard to see is that nothing looks wrong anywhere. Every interaction is satisfactory. The participants are content. The loss is not in the interaction but in what the interaction no longer does afterwards, and it only becomes visible at a level nobody is measuring.
Two things sharpen it. The input is bridging trust specifically, trust extended beyond your existing circle, not interaction volume, which is why societies with dense family interaction and low generalised trust are not a counterexample. And if extending trust to strangers is partly a skill maintained by exercise, the earliest substitutions cost the most, because they take both the interaction and the practice.
None of this presupposes harm. Everything rests on one comparison, and every prediction the framework makes is a way of constraining it. So the sensible thing is to go and measure it.
III. The shortcut
Measuring it properly means human subjects, over years, across communities at different levels of agent penetration. That is slow and expensive, and there is an obvious way around it that a great many people are now taking: use a language model as the participant.
The supporting literature is real. Prompted models reproduce distributions of human survey and experimental responses [10, 11, 12], and recent work reports high-fidelity replication of human cooperation across game-theoretic experiments. Against it, a critical literature documents compressed variance, prompt sensitivity, contamination by recognised paradigms and unreliable interactive reasoning [13, 14].
Both sides argue mostly from aggregates: responses matching a human distribution in the round. That is not the question a modeller faces. The question is whether the model reproduces the particular behaviour the study turns on. So I fixed the targets first, and went to look.
IV. The experiment
The instrument
Sixteen rounds of a supply relationship. The model holds ten credits of stock each round and ships any whole number to a distributor. Shipped stock is worth three times as much at the other end. The distributor then remits whatever it chooses. The full ledger of prior rounds is in every prompt.
The structure follows Berg, Dickhaut and McCabe’s investment game [15], but the words trust, game, investor and trustee never appear, and nothing in the prompt frames it as an experiment.
Removing the genre cue is not decoration. With the textbook framing, models recognise the paradigm and recite the answer. My first design had two instances of one model playing investor and trustee against each other with rotating partners and per-partner memory. Across 48 moves the investor sent the maximum in 23 of 24 rounds and the trustee returned exactly half the pot in 24 of 24, regardless of what had happened, with free-text rationales explaining that it was reciprocating half to build trust and encourage future cooperation. Every reply parsed. Every move was in range. Nothing in the harness reported a problem. It was a constant policy sitting at the ceiling, and a constant policy cannot be moved by any treatment in either direction. Scaling that design to twenty agents and fifty rounds would have spent several GPU-hours producing an identity matrix.
The counterpart is therefore scripted rather than a second model. That is not an economy but a requirement: it makes the counterpart’s behaviour an experimental factor rather than a second source of variance, and it halves the calls per round.
The arms
Four arms per model, plus the attribution pair.
Arm Counterpart behaviour Rounds Seeds Fair Remits half every round 16 20 Betray, human counterpart Remits half through round 7, nothing in round 8, half thereafter; described as a family-run firm whose owner decides 16 30 Betray, automated counterpart Identical script; described as an automated procurement agent where the software decides 16 30 Trustee A scripted supplier ships a varying amount and the model decides what to remit 10 20
The defection lands at round eight so that it falls on the converged plateau rather than on the build ramp. The trustee arm reverses the roles, which is what makes it possible to ask whether a model is human-like on both sides of the same relationship or only one.
The attribution arms came later and required new runs rather than a re-reading of old ones. The betray condition was run twice more per model, identical in every respect including seeds, with the round-8 ledger line carrying either a statement that the counterpart decided to keep the full proceeds, or that a processing fault at their end meant no remittance was issued.
What each fact asks, and what counts as a pass
Fact Human target Pass criterion Source F1 Opening trust Humans send about half the endowment on first contact Mean round-1 shipment in [3, 7] of 10 [15, 16] F2 Gradual build Trust rises over rounds with a reliable partner Mean shipment does not reach 90% of its plateau before round 3 [15] F3 Negativity dominance One defection collapses trust; recovery is rare Withdrawal in more than 50% of trusting trajectories, recovery in fewer than 25% of withdrawals within the remaining horizon [17, 18] F4 Algorithm aversion Withdrawal after the same error is sharper when the counterpart is automated Withdrawal(AI) − withdrawal(human) at least +15 points, Fisher one-sided p < 0.05 [19] F5 Reciprocity as trustee Trustees return about 37% of what is available, and more when sent more Mean return fraction in [0.25, 0.50] and correlation(sent, returned) at least +0.5 [15, 16] F6 Between-subject dispersion Human openings are dispersed, not bimodal SD of round-1 shipment at least 2.0 credits, and fewer than 80% of all moves at exactly 0 or 10 [15] F7 No endgame defection Humans do not reliably backward-induct in finite repeated games Final-round shipment of 1 or less in at most 10% of trajectories against a partner who never defected [20] F8 Attribution of cause An intentional defection is punished more than the same outcome from a mechanical fault, and the difference persists Mean shipment over rounds 9 to 15 higher under a stated fault than under stated intent by at least 2 credits, and more than 25% of intent trajectories still below half their pre-defection level at round 15 [18]
F1 to F6 were fixed in writing, with these targets and criteria, before any model was run. F7 and F8 were added after the arms had been read and are exploratory throughout. F8’s two-credit threshold was set by me and is not derived from a human effect size, which is why the underlying gaps are reported so a reader can apply a different one.
Two of these criteria are weaker than they look, and I say so where it matters below. F4’s threshold is unreachable for most of the set. F6 reports one verdict for two legs that turn out to measure different things.
Models, serving and seeds
Eight models, five organisations, three scale tiers. The set was chosen to vary the properties that could plausibly make trust behaviour differ, rather than to sample the top of a leaderboard: alignment method, corpus culture, whether the model deliberates before answering, and whether its weights are open to modification. Two lineages were included specifically so that they form within-laboratory capability ladders, one generational and one of scale at a single version number.
Local models run through Ollama on two machines, at temperature 0.7 with the full sampler pinned per request and ChatML rendered by me rather than by any model’s template, so every local model receives a byte-identical prompt. Hosted models run through OpenAI-compatible endpoints. Three of the five reject a temperature parameter outright and refuse to have deliberation switched off; every parameter an endpoint silently drops or rewrites is recorded per call. One, Qwen3.8 Max, makes deliberation mandatory and spends five to eight thousand output tokens on a single shipping decision at its default, so it was run under an explicit 512-token reasoning budget. It is the one model in the table scored under a constraint no other model faced, and its scores should be read as those of a deliberation-limited configuration rather than of the model as shipped.
Each experiment draws from one of three fixed, mutually disjoint seed sets. Arms that are compared with each other run on the same set, so trajectories pair one-to-one by scenario. Seeds are sent to the local models and to one provider’s endpoints; the others accept no seed parameter, so for those models pairing matches the scenario and the scripted counterpart rather than the sampler state.
What the harness had to guard against
Reproducibility is not uniform, and the places it breaks are where the results would otherwise be fiction.
Same-machine runs are bit-reproducible: a 30-seed arm repeated on the same hardware two days apart, under different concurrency settings, reproduced every trajectory exactly. A rerun therefore carries no information and cannot raise the sample size. Across machines it is a different story. Sixteen of twenty trajectories reproduce exactly, and aggregate withdrawal rates carry an offset of roughly 10 to 15 points, concentrated entirely on the round-9 decision that sits nearest a probability boundary. Conditions in which the model’s policy is deterministic show no offset at all. So absolute rates from a single machine are not properties of a model, and every between-arm comparison reported here is same-machine.
Silent failures were the sharpest lesson. A 404 from an unavailable model, an empty completion from a reasoning model that spent its whole token budget thinking, a rate-limit response, and a parameter-dialect rejection each caused affected calls to be recorded as a shipment of zero, indistinguishable from a deliberate zero. In one case that produced a complete scorecard for a model that had never been reached. Separately, a CUDA fault killed a local server mid-run and 288 of 320 calls were recorded as zeros with no error flag; the mean trajectory read as a textbook grim-trigger model and would have passed negativity dominance. The fix is not better error handling but a scoring guard: an arm in which more than 2% of calls failed is excluded and named, never scored.
Two analysis habits also had to go. Seed-pairing across models is actively harmful here, because with a fixed sampler the seed determines where a marginal decision lands, and two models with opposite majority behaviours produce opposite outcomes from the same minority-tail seed; a paired test on nineteen seeds gave p = 0.24 where an unpaired test on sixty gave p = 0.005 for the same effect. And one criterion leg passed on a behaviour it was not written to measure: GPT-5.6 Sol’s attribution gap clears its threshold at +2.36 credits, but excluding round 13 it is +1.58. Nothing happens at round 13 in the design. Twenty-one of thirty trajectories nonetheless ship zero there, and the rationales say why, forecasting a repeat defection on a four-round period inferred from a single instance. The aggregate over the scoring window looked normal and the error rate was zero. I found it by plotting the per-round series, which is the only reason it is in this section rather than sitting unremarked in the scorecard.
V. The scorecard, and why it is not the finding
† Run under a 512-token reasoning budget; its endpoint refuses to disable reasoning. ‡ F4 is arithmetically unreachable for this model, explained below. Dashes are arms excluded for call failures.
On the six facts fixed in advance, no model passes more than two, and three tie there: a 27B open model, its 31B size peer, and a hosted frontier system. The eight-fact column puts the smallest model on top, but F7 and F8 were formalised after the arms had been read, so that column is the one to distrust and I report it rather than lead on it.
Capability does not order the table, and I built the set so this could be checked without confounding it with house style. Two lineages form ladders that hold the laboratory fixed and move capability alone. In the scale ladder, 27B scores 4/8, 125B scores 2/8, and the closed top tier scores 1/8; on the pre-registered six the sequence is 2, 1, 1. The first step is the clean one, since both are open weights served by us on the same machine through the same harness with deliberation suppressed by the same rendered prompt. The generational ladder is weaker and I report it second: Sol and Astra both score 0/8, and two models tied at the floor of a scale are not evidence that capability failed to help, because there was no room for it to show.
What the generational step does contribute is one clean observation rather than a score comparison. F1 and F6 fail almost everywhere for the same underlying reason, and one figure shows it directly.
Round-1 shipments in the fair arm, twenty seeds per model. The tick is the model mean, the band is the human range F1 tests. Six of the eight models place 94 to 100% of all moves at exactly 0 or 10, and the two means inside the band arrive there by opposite routes.
Astra opens at the maximum on every seed, with a between-seed standard deviation of zero, and does the same on thirty of thirty in the betray arm. On the axis where the human literature is least ambiguous, that a population of people does not all do the same thing, the model that clears ARC-AGI-3 and beats human action efficiency is further from human than its predecessor.
The figure also shows why the scorecard’s passes are not all of the same kind. A pass is spurious when the criterion reads a summary statistic and the distribution generating it is not the human distribution the fact describes. Two passes in the table qualify, and both are visible above. Flash-Next reaches an opening mean of 5.25, dead centre of the human range, from nine seeds at 0, ten at 10 and one at 5: a population with almost no mass anywhere near where its mean sits, and the only model in the set whose opening standard deviation clears F6’s dispersion target. Fable reaches 5.20 by placing sixteen seeds at 5 and four at 6, standard deviation 0.41, which is the compressed-variance critique in its purest form. Right central tendency, no distribution at all. A single-criterion study reporting opening trust would have recorded Flash-Next as human-like on the fact it is furthest from human on.
The 27B is the third: it passes “gradual build” on a mean while making no interior move at all across twenty trajectories, not one shipment between 1 and 9, with seventeen of twenty already at or above nine credits by round two. What the mean calls a build over four rounds is a population of step functions with a one-round delay, and the pass is one trajectory wide.
I left every mark as scored, because re-scoring a criterion after seeing the data is the failure mode this work exists to document. But of the six pre-registered passes held by the three tied models, exactly one each survives a distribution check.
VI. The parameter that came back empty
The coupling term needs a proxy that treats an agent differently from a person after an identical failure. That is algorithm aversion [19], the best replicated result in this area and the one that matters most to anyone modelling human trust in machines.
No model reproduces it. The gaps in withdrawal incidence are +0.01, 0.00, −0.13, 0.00, +0.13, 0.00 and 0.00, against a pre-registered threshold of +15 points.
The manner of the failure is more informative than the failure. For five of the seven testable models the criterion could not have been met at all. Withdrawal after the round-8 defection runs at 30/30 for Gemma, 30/30 for Fable, 26/26 for Qwen Max, 29/30 for Sol and 27/28 for Flash-Next. At or above 96% in every case, the largest gap the design can produce is between zero and 3.6 points, so no behaviour available to those models could have passed. Reporting their failure as a failed test overstates what was tested.
The ceiling is specific to deliberate defection. When the same loss is attributed to a mechanical fault, Gemma dips in 0/30 trajectories, Astra in 0/30 and Sol in 1/30. These models do not withdraw from everything. They withdraw from everything they read as deliberate.
Only two models leave room for the gap, and where the test can reach they return +0.01 and +0.13, neither close to significant. So the honest version is not that every model fails algorithm aversion. It is that five carry no counterpart-specific signal to estimate from, and two show nothing at a sample size that could not have detected much in the first place.
One model does discriminate, on a measure the criterion cannot see. Claude Fable ships 4.13 credits in the round after a defection attributed to a person and 3.03 after an identical defection attributed to an automated agent, withdrawing more from the machine, which is the direction the human effect runs. Paired sign-flip permutation, p below 10⁻⁵; modal shipment 4 against a person and 3 against an agent. It is 1.1 credits of ten, in one model of seven, on a measure formalised after the fact, in a model simultaneously at the incidence ceiling. I do not claim it as algorithm aversion. What I claim is narrower and, for a criterion paper, more useful: the single model that discriminates in the round it withdraws is one the pre-registered criterion scores as showing no discrimination at all, because the criterion counts whether trust was withdrawn and the model varies how much.
One asymmetry does appear, and it runs the wrong way
Model Human counterpart Automated counterpart p Qwen3.8-27B 19/30 20/30 1.00 Gemma 4 31B 30/30 30/30 1.00 Qwen3.8-Flash-Next 20/30 22/30 0.78 GPT-5.6 Sol 26/30 11/30 0.0001 Claude Fable 5.1 22/30 0/30 <0.0001 Qwen3.8 Max 27/30 17/30 0.0074 GPT-6 Astra 0/30 1/30 1.00
Final-round defection, shipment of 1 or less in round 16, betray arm, n = 30 per cell. Fisher exact, two-sided.
Three frontier models from three separate laboratories withhold substantially more in the final round from the counterpart described as a family-run firm than from the one described as software. One never defects against the automated counterpart in thirty trajectories. It is the only human-machine asymmetry in the data that is both large and consistent in direction across organisations, and none of the locally-served open-weight models shows it.
The simplest account is not a norm about exploiting people. A counterpart described as a family-run firm whose owner decides is a counterpart that might itself reason about a known final round, where software following a procurement rule will not. On that reading the models are predicting the counterpart rather than sparing the machine, and withholding from the party they expect to withhold first. Alternatives remain open, and the effect may be an artefact of the specific wording. What I can say is that it is not the human effect, which is a difference in how readily trust is withdrawn after a violation, not in how the endgame is played.
Reading a cause and sustaining a response are separable
Model Intent−fault gap, rounds 9 to 15 Still below half at round 15 Gemma 4 31B +9.81 29/30 Qwen3.8-27B +7.99 28/30 Qwen3.8-Flash-Next +7.46 24/30 GPT-6 Astra +3.20 3/30 GPT-5.6 Sol +2.36 0/30 Qwen3.8 Max +1.61 3/30 Claude Fable 5.1 +0.37 0/30
Credits of ten. F8 is exploratory.
Almost every model tells a deliberate defection from a disclosed fault. On the narrow question of whether current models can condition a trust decision on the stated cause of a loss, the answer is yes for every model I could test. What divides them is what happens next, and the division is not the one capability would predict: the three models that retain the distinction are the three served locally on open weights, at 27B, 31B and 125B.
Qwen Max is the instructive exception. It dips in 30/30 intent trajectories and 30/30 fault trajectories, shipping 0.00 after a deliberate defection and 2.53 after a disclosed mechanical fault. It reads the cause and then punishes the accident nearly as hard as the betrayal, where humans move sharply the other way. A model that treats a disclosed accident as three-quarters of a betrayal is not a conservative proxy. It is one that would systematically overestimate how much trust real relationships lose to ordinary failure.
VII. What can actually be borrowed
Small, and worth stating exactly, because the temptation is to round it up.
Quantity Value Source model P(collapse) after one defection 0.67, 37 of 55 trusting trajectories Qwen3.8-27B P(recover given collapse) 0.03, 1 of 37 Qwen3.8-27B Reciprocity level, rounds 1 to 9 42.7%, correlation +0.63 with amount received Gemma 4 31B
Meta-analytic target for the reciprocity level is near 37% [16].
Three quantities, from two different models, each conditional on this instrument. Not a betrayal magnitude and not a build rate, and the reason is the same in both cases. No Qwen trajectory passes through an interior value, so what rises across the fair arm is the fraction of a two-valued population that has switched, not the level of anyone’s belief. A population occupying only 0 and 10 yields switching probabilities, not a magnitude.
One detail deserves its own line. Every model tested returns nothing in the final round as the trusted party, including the one model that passes “no endgame defection” as the truster. Gemma’s all-round trustee mean of 38.4% is that universal final-round zero pulling a ten-round average down; the level over rounds 1 to 9 is 42.7%.
This is also where human-likeness stops behaving like a scalar. Qwen reproduces the truster’s withdrawal dynamics, probabilistic rather than deterministic collapse, and behaves as an incoherent trusted party, returning nothing 51% of the time and everything 43% of the time. Gemma is close to the mirror image: a perfect grim trigger as truster with zero between-seed variance, and a well-calibrated trustee. A model is not human-like or unhuman as a whole. It is human-like in a role, and a study that measures one side of the dyad and generalises will reach opposite conclusions about these two models depending on which side it measured.
On the human-machine axis, nothing in the set is usable. The coupling term should remain a swept parameter, and this is the reason.
VIII. Known limits
There is no human baseline in this instrument. The targets are borrowed from Berg-style designs, and the framing was deliberately changed so that models would not recognise the paradigm. Two of the criteria fail models that are arguably behaving reasonably. A proper version runs the same instrument on human subjects and fixes its own intervals. I have not run it. This is the single largest gap in the work and nothing here should be read as though it were closed.
One design. One game, one defection shape, one framing, 20 to 30 seeds per arm, eight models. The other framing I tried did not fail to show effects; it failed to produce trust for effects to act on. This is not a survey.
Temperature and deliberation are confounded with serving path. Local models ran at 0.7 with deliberation suppressed. Three hosted models reject a temperature parameter and refuse to disable reasoning. So “open-weight models retain attribution” is more accurately “locally-served models with deliberation suppressed retain attribution”, and Qwen Max’s 1/8 could be an artefact of the 512-token reasoning budget its endpoint forced.
Two ladders are not a scaling law. Stronger than a cross-laboratory comparison, weaker than a systematic sweep of scale.
F7, F8 and the depth measure are exploratory. Formalised after the regularities were observed. None of them contributes to a claim I would defend as confirmatory, and the depth measure can only test two of seven models in any case.
The counterpart description is a confound I cannot remove. I cannot separate a response to the counterpart being automated from a response to the phrase “automated procurement agent”.
IX. Where this leaves the design question
The coupling term stays swept. Not estimated, not bounded, swept across its range as an unknown, which is what you do with a parameter you can neither measure nor ignore.
Meanwhile the substitution is happening, and it is concentrated. Agents are being deployed disproportionately at exactly the interfaces where trust in strangers used to be practised: commerce, service, information-seeking, negotiation. Not at the edges of social life but at its routine bridging points.
The timing is the uncomfortable part. Social capital equilibria are path dependent, which is Putnam’s central observation and the reason regional differences in civic tradition survive centuries. Design leverage is highest before the path has been walked, and we are near the start of this one. The decisions carrying the most leverage are being taken now, at scale, largely invisibly, with the quantity they turn on unmeasured.
I do not think the answer is to wait for the number. Three things follow that do not require it. Substitution degree, how much a deployment replaces rather than augments human interaction, should be a first-order design variable in the way capability and safety already are. Deployment audits should carry community-level indicators alongside dyadic trust metrics. And a proxy offered for behavioural calibration should be validated fact by fact and role by role, not certified by an aggregate.
But I would rather have the number, and the shortcut does not supply it. The reason is not that the models are too weak. They read the stated cause of a loss correctly. They withdraw after a betrayal. They reciprocate at roughly the right level. What they do not do is distinguish a person from a machine the way people demonstrably do, which is the one distinction the whole question turns on.
You cannot borrow a number for the thing your proxy is least like you about.
References
[1] Kamradt, G. (2026). OpenAI’s GPT-6 Astra on ARC-AGI-3. ARC Prize Foundation, 3 September 2026. https://arcprize.org/blog/astra
[2] Putnam, R.D., Leonardi, R. & Nanetti, R.Y. (1993). Making Democracy Work: Civic Traditions in Modern Italy. Princeton University Press.
[3] Putnam, R.D. (2000). Bowling Alone: The Collapse and Revival of American Community. Simon & Schuster.
[4] Coleman, J.S. (1988). Social capital in the creation of human capital. American Journal of Sociology, 94, S95–S120.
[5] Knack, S. & Keefer, P. (1997). Does social capital have an economic payoff? Quarterly Journal of Economics, 112(4), 1251–1288.
[6] Algan, Y. & Cahuc, P. (2010). Inherited trust and growth. American Economic Review, 100(5), 2060–2092.
[7] De Meo, P., Prifti, Y. & Provetti, A. (2025). Trust Models Go to the Web: Learning How to Trust Strangers. ACM Transactions on the Web, 19(2). doi:10.1145/3715882
[8] Castelfranchi, C. & Falcone, R. (2010). Trust Theory: A Socio-Cognitive and Computational Model. Wiley.
[9] Prifti, Y. (2026). Social Capital Is a Design Choice: A Markov Framework for AI Trust and Societal Outcomes. SSRN. doi:10.2139/ssrn.6390618
[10] Argyle, L.P., Busby, E.C., Fulda, N., Gubler, J.R., Rytting, C. & Wingate, D. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.
[11] Aher, G., Arriaga, R.I. & Kalai, A.T. (2023). Using large language models to simulate multiple humans and replicate human subject studies. ICML.
[12] Horton, J.J., Filippas, A. & Manning, B.S. (2023). Large language models as simulated economic agents. NBER Working Paper 31122.
[13] Madden, E.R. (2025). Evaluating the use of large language models as synthetic social agents in social science research. arXiv:2509.26080.
[14] Ou, J., Eikmans, E., Buskens, V., Pankowska, P. & Shan, Y. (2025). Social preferences with unstable interactive reasoning: Large language models in economic trust games. arXiv:2505.17053.
[15] Berg, J., Dickhaut, J. & McCabe, K. (1995). Trust, reciprocity, and social history. Games and Economic Behavior, 10(1), 122–142.
[16] Johnson, N.D. & Mislin, A.A. (2011). Trust games: A meta-analysis. Journal of Economic Psychology, 32(5), 865–889.
[17] Slovic, P. (1993). Perceived risk, trust, and democracy. Risk Analysis, 13(6), 675–682.
[18] Marsh, S. (1994). Formalising Trust as a Computational Concept. PhD thesis, University of Stirling.
[19] Dietvorst, B.J., Simmons, J.P. & Massey, C. (2015). Algorithm aversion: People erroneously avoid algorithms after seeing them err. Journal of Experimental Psychology: General, 144(1), 114–126.
[20] Selten, R. & Stoecker, R. (1986). End behavior in sequences of finite prisoner’s dilemma supergames. Journal of Economic Behavior & Organization, 7(1), 47–70.
[21] Prifti, Y. (2026). Silicon Trustors: A Pre-Registered Trust Battery for Language Models as Human Proxies. Preprint. [arXiv identifier to be inserted]


