No synthetic-research vendor publishes the science against their own model. We do, because we’re named for it. A deep-research run pulled 110 claims from 22 sources; the top 25 went through adversarial verification — independent 3-voter skeptic panels arguing each claim before it could stand. All 25 survived 3–0.
104agents ran the research
22sources fetched and read
110 → 25claims extracted, then selected for adversarial review
25 / 25survived independent 3-voter panels, 3–0
Where it fails
Five ways synthetic respondents deceive — verified, not asserted.
Each finding below is HIGH confidence: independently replicated, peer-reviewed, and adversarially verified against our own skeptic panel.
GPT–human response similarity falls as cultural distance from the United States grows.
r = −.70
(ρ = −.72, robust to multilevel checks.) Measured across the World Values Survey — 94,278 humans, 65 nations, 262 variables. In hierarchical clustering GPT sits closest to the US and Uruguay, farthest from Ethiopia, Pakistan, and Kyrgyzstan. Independently replicated across five GPT generations and 107 countries; partially mitigable by cultural prompting, not solved by scale.
Atari, Xue, Park, Blasi & Henrich 2023, “Which Humans?” (PsyArXiv/Harvard) · Tao et al. 2024, PNAS Nexus (arXiv:2311.14096)
“WEIRD in, WEIRD out.”
Every synthetic mean fell within one standard deviation of the human average. Underneath the mean, the structure disagrees.
48% · 32%
48% of regression coefficients differed significantly from human data — and the sign flipped in about 32% of those cases. Separately, LLM respondents produced significant results for up to 83% of effects that were null in humans, even at an overall correlation of r ≈ 0.85. The most dangerous failure mode: plausible-looking data leading to wrong decisions.
Bisbee et al. 2024, Political Analysis 32(4), doi:10.1017/pan.2024.5 · Cui et al. 2025, Nature Computational Science (s43588-025-00840-7), via Hullman et al. arXiv:2602.15785
SD 16.1 vs 31.4
Variance compression
ChatGPT-simulated ANES respondents showed roughly half the standard deviation of real humans — badly enough that a power analysis on synthetic data would call for ~33 respondents where real variance requires nearly an order of magnitude more. Shown to be structural to likelihood training, demonstrated across 4 LLMs with 3,200 human participants across 16 demographic identities.
Bisbee et al. 2024, Political Analysis 32(4) · Wang, Morgenstern & Dickerson 2025, Nature Machine Intelligence (s42256-025-00986-z)
Not reproducible
Prompt-brittle, model-brittle
Identical prompts run in April, June, and July 2023 gave different distributions after an OpenAI model update — “neither recollection of our data exactly reproduces our original synthetic data.” Switching prompt language (English ↔ Chinese) or assigned role significantly changed behavior in an identical economic game. A research product built on third-party models inherits silent breakage of longitudinal comparability.
Demographic and partisan caricature — worst for minorities
Simulated partisans voted their own party at 99.96% / 99.22% — real humans don’t. Out-group antipathy ran 10–20 thermometer points more extreme than real Democrats; LLM responses resembled out-group stereotypes of a group more than the group’s own self-reports. Real white/Black respondents rated racial diversity “better” at 54%/55%; synthetic personas said 87%/98%. Main-effect replication drops 77%→42% on socially sensitive topics — RLHF harmlessness training distorts minority-group personas hardest.
Bisbee et al. 2024 · Sun et al. 2024, “Random Silicon Sampling” (arXiv:2402.18144, LREC-COLING) · Wang et al. 2025 Nature MI · Cui et al. 2025 Nature CS
p < 0.001
Behavioral simulation fails even in simple games
All 8 tested models (GPT-4/3.5, Claude 3 Opus/Sonnet, Llama 2/3) diverged from human distributions in the 11-20 money-request game, clustering at level-0/1 strategic reasoning against a human level-3. Chain-of-thought, personas, few-shot, and RAG all failed to close the gap; only fine-tuning on the game’s own human data worked — and it didn’t generalize. Newer reasoning models overshoot toward Nash, deviating in the opposite direction.
Gao et al. 2025, PNAS · corroborated arXiv:2404.08492
Where it works
A narrow, honest niche — and it’s real.
The synthetic layer earns its place in early screening: cheap, fast, and directionally useful, inside a ceiling we state plainly.
MEDIUM confidence — single preprint, one 2023-era model
Cheap exploratory piloting is where synthetic panels are genuinely extraordinary.
5,441 responses · $0.70
Collected in about an hour. Useful for ideation, questionnaire debugging, and directional concept ranking (effect-direction correlation with humans r ≈ 0.85). But replicability breaks below ~200 samples, and none of it is verified outside a 70%-Caucasian US population.
Sun et al. 2024, arXiv:2402.18144
KL 0.0004
Population-level aggregates, on non-sensitive questions
GPT-3.5 reproduced the 2020 Biden/Trump vote split at a KL-divergence of 0.0004 (57.43% ± 0.46% vs. 58.88% actual) — but statistically replicated only 1 of 10 other ANES opinion topics, with extreme response tendencies on six of them. One result, one model, and a real memorization risk (the vote outcome may have been in training data).
Sun et al. 2024, arXiv:2402.18144
The augmentation case
Human-in-the-loop is supported — with a refinement that changes the roadmap.
“Always keep a human in the loop” is not the winning claim the industry assumes. The evidence says something more precise.
Statistically calibrating LLM data against a human sample — the principled version of the vendor pitch — still runs into a hard ceiling.
+13% effective n
Adding 100,000 LLM responses to 10,000 human ones (prediction-powered inference) yielded only about a 13% effective-sample-size gain — 10,000 humans grew to at most 11,275 effective respondents. The heuristic “validate-then-simulate” approach vendors use (show correlation with a human benchmark, then substitute) cannot guarantee the absence of systematic bias, and is unsuitable for confirmatory research. Five independent author teams across five disciplines converge on the same word: supplement, never replace — even Qualtrics, commercially motivated to say otherwise, writes that “synthetic data augments human research, it does not replace it… synthetic models degrade without ongoing human data collection.”
Hullman, Broska, Sun & Shaw 2026, arXiv:2602.15785 · Broska, Howes & van Loon 2025, Sociological Methods & Research (doi:10.1177/00491241251326865)
g = −0.23
The refinement: human–AI combos don’t always win
A meta-analysis of 106 experiments and 370 effect sizes found human–AI combinations performed, on average, worse than the best single agent alone (Hedges’ g = −0.23, CI −0.39 to −0.07). It’s task-dependent: decision tasks negative (g = −0.27), content-creation positive (g = +0.19). Direction depends on who’s stronger — human stronger, combo wins (g = +0.46); AI stronger, combo loses (g = −0.54). Combos do reliably beat unaided humans (g = +0.64). Explanations, confidence displays, and division-of-labor were not significant moderators — task type was.
MIT meta-analysis, Nature Human Behaviour 2024 (s41562-024-02024-1)
The product implication: “put the human where the human is superior” — research design, cultural interpretation, field access in low-data markets, judgment on contested findings. This is the scientific basis for our expert routing layer.
Industry signals
Extracted, but not yet put through the panel.
The verification method cuts both ways: it also tells you what we haven’t verified. These are vendor and industry sources, treated as leads, not as claims we’d stand behind. Absence of surviving vendor evidence isn’t proof vendor evidence is wrong — the corpus simply skews academic.
Not adversarially verified. Included for completeness and transparency about what our own review has and hasn’t tested.
PyMC Labs × Colgate-Palmolive
Semantic Similarity Rating reached ~90% of human product-ranking correlation attainment against 57 real surveys (9,300 responses), >85% distributional similarity. Vendor preprint (arXiv:2510.08338) with open-source code; not peer-reviewed, no limitations discussed.
Kantar
Ran GPT-4 against ~5,000 human respondents on Likert and open-ended questions. Public stance: hybrid augmentation (data quality, fraud detection, translation), not synthetic replacement.
NIQ
“Synthetic respondents are not a replacement for human consumers… a supplement to your ideation process when time is of the essence.”
LISS panel ground-truth test
Digital personas improved aggregate distributional alignment but failed at individual-level prediction and multivariate structure (arXiv:2605.10659) — consistent with the academic failure modes above.
Two honesty notes
Model vintage: most of the damning numbers are GPT-3.5/4-era (2023–24). Late-2025/26 replications found the biases persist or shift rather than disappear — Nature MI argues the flattening is structural. Still, re-check against 2026 frontier-model studies before any investment decision.
One-sided survivorship: no vendor validation study survived 3-vote adversarial verification; the fetched corpus skews academic and critical. Absence of surviving vendor evidence is not evidence that vendors are wrong.
This is why we route, not replace.
Every study we run gets sent to the mode the science supports — synthetic screening where data is rich, human fieldwork where it isn’t, an expert deciding which.