EvidenceField
A woven fiber sphere — the human, field-collected material the argument turns on.

A field manifesto

Research is only as global as its evidence.

“WEIRD in, WEIRD out.”


01  /  The premise

The tooling assumes a world that isn’t there.

Zoom automation, web panels, transcription pipelines — the modern research stack quietly assumes a digitally fluent participant. In Vietnam, Indonesia, Brazil, Mexico, people often won’t, or can’t, complete digital onboarding or sit for a remote study.

We learned this operating a marketplace across 130+ countries: the method breaks long before the insight does.


02  /  The failure mode

Synthetic respondents fail exactly where research is hardest.

GPT–human response similarity falls with cultural distance from the United States at r = −.70.Atari et al. 2023, World Values Survey

The pattern replicated across five GPT generations and 107 countries.Tao et al. 2024, PNAS Nexus

The model is most confident where it is least valid. WEIRD in, WEIRD out.


03  /  The distortion

The averages look right. The relationships are wrong.

Synthetic panels compress variance to about half of real humans — SD 16.1 against 31.4. Roughly 48% of regression coefficients differ significantly from human data, and the sign flips in about 32% of those.

Identical prompts on different model versions return different distributions. RLHF caricatures minority groups worst of all. The mean survives; the structure does not.


04  /  The honest case

And yet — used honestly, it earns its place.

For cheap exploratory piloting, ideation, and questionnaire debugging, synthetic panels are extraordinary: about 5,441 responses for roughly $0.70, inside an hour.

The discipline is knowing the ceiling. Replicability breaks below ~200 samples, and none of it is verified outside WEIRD populations. Use it where being wrong is cheap.


05  /  The consensus

Everyone serious has already converged.

Kantar, NIQ, Qualtrics — the incumbents ran the experiments and retreated to the same line.

“Synthetic data augments human research, it does not replace it.”
— Qualtrics

06  /  The scarce asset

Human data is the scarce asset — not an inefficiency.

Even the best statistical hybrid — calibrating 100,000 model responses against 10,000human ones — adds only about 13% effective sample size.

The bottleneck is real people, in hard places, answering honestly. Whoever can collect real human data cheaply in difficult markets owns the one thing a model upgrade can’t erode.


07  /  The division of labor

So put each in its lane.

The MIT and Nature Human Behaviour meta-analysis of 106experiments is blunt: human–AI teams win only when the human is the stronger performer, and lose when the AI is.

“Put the human where the human is superior.”

The creed

Synthetic where data is rich.
Human where it isn’t.
Expert judgment deciding which.

Read the plan