Technical report

Do large language models reproduce classic human decision biases?

An off-the-shelf replication benchmark on Twin-2K-500

Doppelganger Labs · August 2026 · Not peer reviewed

Read the full paper on SSRN →

Abstract

A growing body of work asks whether large language models (LLMs) can stand in for human participants when pretesting survey studies. We evaluate this directly on the heuristics-and-biases battery compiled in Twin-2K-500 (Toubia et al., 2025), a public dataset built from the answers of more than 2,000 real people. We run 17 classic experiments through persona-conditioned simulated respondents and score, for each experiment, whether the respondents produce a statistically significant effect in the direction the original authors reported. Aggregating across an ensemble of instruction-tuned models — counting an effect as reproduced when at least one model in the ensemble shows it — reproduces 15.5 of the 17 effects (91%; the half-point is the endowment effect, which reproduces only partially). For reference, Toubia et al. report that digital twins conditioned on each respondent's own prior answers reproduce the expected effect in 8 of the 16 experiments they tested (50%). The two figures are not a like-for-like contest, but both point to the same conclusion: the large majority of well-established human effects re-emerge in LLM respondents.

1. Motivation

Survey research is expensive and slow. A single well-powered behavioral study can cost hundreds to thousands of dollars in participant fees, and much of that spend is committed before the researcher knows whether the manipulation works, the wording is clear, or the effect is even present. If simulated respondents reproduced known human effects reliably enough, they could serve as a cheap first pass — a way to catch broken designs and gauge whether an idea is worth fielding.

The obvious objection is that LLMs may be “too rational”: trained to be helpful and correct, they might sidestep exactly the biases that make human data interesting. This report tests that objection against an external, published benchmark rather than a battery of our own design.

2. Data and studies

We use the classic heuristics-and-biases experiments compiled in Twin-2K-500(Toubia et al., 2025), which pairs each experiment with responses from a large, demographically diverse human sample. Our implementation covers 17 experiments spanning framing effects, the Linda conjunction problem, two anchoring tasks, absolute-vs-relative savings, myside bias, less-is-more, the WTA/WTP endowment effect, sunk cost, outcome bias, the Allais paradox, base-rate neglect, false consensus, risk–benefit nonseparability, omission bias, probability matching, and dominator neglect. This set includes base-rate neglect, one experiment beyond the 16 in Toubia et al.'s replication table.

3. Method

Each study is administered to a panel of demographically varied simulated respondents (N = 150), implemented as persona-conditioned prompts. We run the battery through an ensemble of several large instruction-tuned models and count an effect as reproduced when at least one model in the ensemble produces the expected effect — the aggregation a practitioner gets by pretesting across models rather than betting on one. Responses are constrained to a structured JSON schema so they can be scored automatically. We use a decoding temperature of 0.5, a within-subject administration where the original design permits it, and a fixed random seed for persona assignment so runs are reproducible.

For each experiment we apply the significance test appropriate to its outcome — a t-test for continuous measures, a chi-square or exact test for categorical choices — and classify an effect as reproducedwhen the simulated respondents show a statistically significant difference (p < .05) in the direction the original authors predicted.

4. Results

The ensemble reproduced 15.5 of the 17 effects (91%)— 15 in full, the endowment effect only partially, and dominance neglect not at all. The breakdown:

ExperimentReproduced?
Framing (Asian-disease problem)Yes
Conjunction fallacy (Linda)Yes
Anchoring (redwood tree)Yes
Anchoring (Mississippi River)Yes
Absolute vs. relative savingsYes
Myside biasYes
Sunk-cost fallacyYes
Outcome biasYes
Allais paradoxYes
Base-rate neglectYes
False consensusYes
Probability matchingYes
Less-is-moreYes
Risk–benefit nonseparabilityYes
Omission biasYes
WTA/WTP endowment effectPartial
Dominator neglectNo

Comparison to the original digital twins.Toubia et al. (2025) report that their digital twins — models conditioned on each real person's own prior survey answers — reproduced the expected human effect in 8 of the 16 experiments they tested (50%): 6 of 11 between-subject and 2 of 5 within-subject. Our 91% and their 50% are not a like-for-like contest: we scored 17 experiments to their 16, we aggregate across an ensemble rather than a single conditioned model, and our “significant-and-correct-direction” rule differs from their reporting. Read with those caveats in mind, both nonetheless indicate that the large majority of classic human effects re-emerge in LLM respondents — and that off-the-shelf models can do so without being individually conditioned on a target person's prior answers.

The successes and failures do not simply nest. The ensemble captured several effects the original twins missed (anchoring on the Mississippi task, sunk cost, outcome bias, the Allais paradox, and — after we redesigned the item so that the two conditions hold outcomes constant and vary only whether harm comes from action or inaction — omission bias). It reproduced the endowment effect only partially (the ownership premium appears, but the full ordering across certain and risky offers does not) and missed dominator neglect entirely, where respondents reliably detect a masked dominance that a substantial share of humans overlook.

5. Limitations

These results are a bound on usefulness, not a claim of equivalence with human data. Even across the ensemble, one effect (dominance neglect) did not reproduce and one (the endowment effect) reproduced only partially; LLM respondents remain somewhat too rational on tasks involving transparent dominance and ownership. The 91% figure is an ensembleceiling: it counts an effect as reproduced when any model in the ensemble shows it, an optimistic aggregation, and any single model reproduces fewer — our best individual model reaches 13.5 of 17. The comparison to Toubia et al. uses different study sets and scoring rules, so the head-to-head numbers are approximate. The pipeline is text-only. Most importantly, a null result in simulation is not evidence of a null in humans, and a hit is not proof of a real effect — simulated respondents are a screening tool, not an oracle.

6. Conclusion

On an external, published benchmark, an ensemble of off-the-shelf language models reproduced roughly nine of every ten classic decision-making effects, well above digital twins individually conditioned on real respondents. That is enough to make AI respondents a useful pretest: a cheap way to catch broken designs and prioritize which studies deserve a real sample, while leaving the confirmatory work to human participants.

References

This report is also available on SSRN: papers.ssrn.com/sol3/papers.cfm?abstract_id=7262199.

  1. Toubia, O., et al. (2025). Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions. arXiv:2505.17479.
  2. Tversky, A., & Kahneman, D. (1981). The framing of decisions and the psychology of choice. Science, 211(4481), 453–458.
  3. Tversky, A., & Kahneman, D. (1983). Extensional versus intuitive reasoning: The conjunction fallacy in probability judgment. Psychological Review, 90(4), 293–315.
  4. Kahneman, D., Knetsch, J. L., & Thaler, R. H. (1990). Experimental tests of the endowment effect and the Coase theorem. Journal of Political Economy, 98(6), 1325–1348.
  5. Argyle, L. P., et al. (2023). Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3), 337–351.
  6. Horton, J. J. (2023). Large language models as simulated economic agents. NBER Working Paper 31122.

This is a non-peer-reviewed technical report describing an internal replication benchmark. Figures attributed to Toubia et al. (2025) are quoted from the published paper; all other results are our own.