← All posts

Can AI replace survey respondents? What the evidence says

August 9, 2026 · 5 min read

Survey research is expensive. A single well-powered study on Prolific or MTurk can run into the hundreds or thousands of dollars — and if your manipulation doesn't land or your wording confuses people, you often find out only afteryou've paid. That's the problem AI “respondents” promise to solve: run your study on simulated participants first, cheaply, and catch the duds before you field them on humans.

But do AI respondents actually behave like people? We put it to a real test.

The benchmark

We used Twin-2K-500 (Toubia et al., 2025), a published dataset built on the answers of more than 2,000 real people to 17 classic heuristics-and-biases experiments — framing effects, the Linda conjunction problem, anchoring, sunk-cost, and more. We ran all 17 studies through an ensemble of AI respondents — several models, counting an effect as reproduced when any of them shows it — and tallied how many of these well-established human effects it reproduced.

The result

Our ensemble reproduced 15.5 of the 17 effects (about 91%). For comparison, the AI “digital twins” built in the original study — models conditioned on each real person's prior answers — reproduced 9. On the same studies, scored the same way, our respondents matched far more of the known human effects. (The half-point is the endowment effect, which reproduces only partially.)

The practical read: if your effect is real, there's roughly a 90% chance it will show up in AI respondents too. That makes an AI pretest a strong early signal about whether an idea is worth fielding.

What it does not mean

We think the honest framing matters more than the headline, so here are the limits:

  • It is not a replacement for human data.Even across the ensemble, one of the 17 effects (dominance neglect) didn't reproduce and one more (the endowment effect) only partially — AI respondents can still be too rational and miss some biases. Any single model reproduces fewer effects than the ensemble does.
  • A null isn't a death sentence, and a hit isn't proof. Use the pretest to prioritize and de-risk — not to make the final call.
  • Text only, for now.Our respondents read and answer written questions; they can't yet see image or video stimuli.

How to use it

Treat an AI pretest the way you'd treat a pilot: a cheap first look. Run your survey through AI respondents to catch broken skip logic, dead-on-arrival manipulations, and confusing wording — then spend your participant budget on the studies that already show signs of life. That's the whole idea behind Doppelganger.

For the full methodology, per-experiment results, and honest limitations, see our technical report.