EvidenceField
Private draft

Business Plan · Working draft · July 2026

EvidenceField

Synthetic AI respondents where data is abundant. AI-augmented human fieldwork where it isn’t. Expert research design deciding which.

Abstract globe, half translucent glass and half woven natural fiber, merging at a seam of light

Research tooling was built for the West. The next billion users weren’t.

The problem

Remote panels, Zoom automation, web-form recruiting, and transcription pipelines all assume digitally fluent, English-adjacent participants. In Vietnam, Indonesia, Brazil, Mexico, and most of the markets where the next billion users live, participants often won’t — or can’t — do digital onboarding or remote studies. Global product teams either fly researchers in at enormous cost, or skip the research and guess.

The false fix

The AI industry’s answer is synthetic respondents: LLM personas that stand in for real participants. It demos beautifully — and the peer-reviewed evidence (Section 03) shows it is least valid exactly where research is hardest. GPT’s similarity to real humans falls with a country’s cultural distance from the United States at r = −.70. Pure-synthetic vendors accelerate research where it was already easy, and fail where the need is greatest.

The model — market-adaptive research

Data-rich markets

Synthetic-heavy

Persona panels for early screening, concept ranking, and questionnaire debugging. Real humans confirm; AI explores.

Data-thin markets

Human-heavy

In-person and low-tech recruitment and fieldwork. AI transcribes, translates, and synthesizes the multilingual qual.

Every market

Expert routing

Research design, mode routing, cultural interpretation, and judgment calls — the human layer that decides what runs where.

Why now

  • AI is finally good enough at the right jobs — transcription, translation, and cross-market synthesis of messy multilingual qualitative data — to collapse the cost of processing human research, without pretending to replace the humans.
  • The failure modes of pure-synthetic research are now documented in top journals — buyers are about to get burned, and the honest alternative doesn’t exist yet.
  • Product organizations are more global than their research stacks. The gap grows every quarter.

Founder edge

Eleven-plus years operating a $2B+ earner marketplace across 130+ countries: firsthand operating knowledge of why Western research and onboarding playbooks fail in Southeast Asia and Latin America, an operator’s network in those markets, and the product + operations discipline that the expert routing layer productizes.

Everyone sells the synthetic layer. Nobody owns the routing — or the field.

Pure-synthetic vendors

Fantasy’s Synthetic Humans (the category’s flagship — Fast Company World Changing Ideas 2024; BP, LG, Ford, Spotify, Google engagements), Synthetic Users, Evidenza, Yabble. Common shape: a persona prompt layer on third-party foundation models. Thin moat — every model upgrade erodes last quarter’s tuning — and structurally weakest in non-WEIRD markets. Their documented failure modes are our “why us.”

Research incumbents

Kantar, NIQ, and Qualtrics have all publicly landed on “supplement, not replace” — NIQ: “synthetic respondents are not a replacement for human consumers… a supplement to your ideation process”; Qualtrics:“synthetic data augments human research, it does not replace it… models degrade without ongoing human data collection.” They validate the hybrid thesis but carry panel-era cost structures and lack emerging-market, low-tech field operations DNA.

The gap nobody covers

Emerging-market fieldwork operations + AI synthesis + expert routing, sold as one system. Pure-synthetic vendors structurally can’t serve non-WEIRD markets; incumbents have panels but not in-person operational muscle in SEA/LatAm; nobody productizes the routing decision itself.

The moat

The human fieldwork network and the routing expertise. Model upgrades erode the synthetic layer — which we deliberately treat as a commodity input — and strengthen our synthesis layer. The proprietary asset compounds: every study adds verified human data from markets where nobody else can collect it.

25 claims, adversarially verified. The science picked this model — we just built for it.

Independent deep-research pass, July 2026: 22 sources fetched, 110 claims extracted, top 25 verified by three-voter skeptic panels — all 25 survived. Venues: Political Analysis, PNAS, Nature Machine Intelligence, Nature Computational Science, Nature Human Behaviour.

Five verified failure modes of synthetic respondents

1

Variance compression

Simulated respondents show roughly half the standard deviation of real humans (16.1 vs 31.4 on ANES data). The flattening is structural to how LLMs are trained.

Bisbee et al. 2024, Political Analysis · Wang et al. 2025, Nature Machine Intelligence
2

Means right, relationships wrong

48% of regression coefficients diverge from human-derived ones; signs flip in 32% of those cases. Up to 83% of human-null effects come back “significant” — even at r ≈ 0.85 overall correlation.

Bisbee et al. 2024 · Cui et al. 2025, Nature Computational Science
3

Non-reproducible

Identical prompts months apart returned different distributions after a vendor model update; changing prompt language changed behavior in an identical task.

Bisbee et al. 2024 · Gao et al. 2025, PNAS
4

Demographic caricature

Simulated partisans vote their party at 99%+; personas resemble out-group stereotypes more than the group’s own self-reports; safety training distorts minority-group personas hardest.

Bisbee 2024 · Sun et al. 2024 · Wang 2025 · Cui 2025
5

WEIRD bias

Across 94,278 humans in 65 nations, GPT-human similarity falls with cultural distance from the US at r = −.70 — replicated across five GPT generations. “WEIRD in, WEIRD out.”

Atari et al. 2023 (Harvard) · Tao et al. 2024, PNAS Nexus

Where synthetic genuinely works

Cheap, fast exploratory work: ideation, questionnaire debugging, directional concept ranking (~5,400 responses for under $1 in an hour; effect-direction correlation with humans r ≈ 0.85). Always a precursor to human data, never a substitute. Better elicitation methods (e.g. PyMC Labs × Colgate-Palmolive’s semantic similarity rating, ~90% ranking attainment) keep improving this layer — which is why we treat it as a commodity input, not a moat.

Why humans stay in the loop — precisely

Human data is the binding constraint: even the best statistical hybrid — calibrating 100,000 LLM responses against 10,000 human ones — adds only ~13% effective sample size (Broska et al. 2025, Sociological Methods & Research). And the MIT meta-analysis of 106 human-AI experiments (Nature Human Behaviour 2024)shows combinations win when the human is the stronger performer and lose when the AI is. So the design rule is not “always add a human” — it is put the human where the human is superior: research design, field access, cultural interpretation, judgment. That rule is the product.

Services first. Product second. The network is the prize.

P0

Foundation — now

  • Resolve name, brand, entity structure, and employment/IP clearance before anything goes public.
  • Use this site to pressure-test the thesis with friends, operators, researchers, and potential partners; recruit a small advisor circle (research methodology + one in-market operator per launch country).
  • Define the wedge study offer and pricing hypothesis.
P1

Expert-led studies — the services wedge

  • Sign 3–5 design partners: tech companies entering or scaling in SEA/LatAm (marketplace, fintech, gig — categories where the founder’s operating credibility is strongest).
  • Run paid, expert-designed studies end-to-end: synthetic pretest where valid, human fieldwork through local recruiters and moderators in Vietnam, Indonesia, Brazil, Mexico; AI transcription, translation, synthesis; decision-ready readouts with a full citation trail.
  • Revenue from day one; every study builds the field network and the evidence corpus. Success gate: repeat purchase from ≥2 design partners.
P2

Productize the system

  • Routing engine: codify the market/data-richness assessment and study-design decisions into software an expert operates — expert in the loop, not expert as bottleneck.
  • Synthesis pipeline: multilingual transcription → translation → theme extraction → evidence ledger, as a repeatable product.
  • Client surface: study tracker + evidence-linked insight library replacing the readout deck.
P3

Scale the network

  • Field network becomes a standing panel asset in markets nobody else covers — refreshed by ongoing studies, feeding calibration data no competitor has.
  • Explore fine-tuning market-specific models on proprietary human data (the one approach that has passed behavioral-fidelity tests — validity to be re-proven per market).
  • Business model shifts from services to platform: subscription + per-study, with services retained for the expert layer.

Risks & honest answers

  • Frontier models close the validity gaps.Possible on variance and elicitation; the WEIRD data gap requires data that doesn’t exist online — which our field network generates. We commoditize the layer that improves and own the layer that doesn’t.
  • Ops-heavy business. Yes — deliberately. Start narrow (2 markets, 1 vertical), let AI absorb the processing cost, and price as research value, not hours.
  • Buyer skepticism of anything “synthetic.” Our positioning is the skepticism, with citations. We sell the honest version of the category.

Not interviews. Not a dashboard. Defensible decisions in markets you can’t see.

Three layers

Layer one

Synthetic Panel Studio

Persona panels on commodity foundation models for exploratory screening in data-rich markets: concept ranking, questionnaire debugging, hypothesis generation. Clearly labeled exploratory; never confirmatory.

Layer two

Field Network

Vetted local recruiters, moderators, and translators in data-thin markets, running in-person and low-tech studies — voice notes, WhatsApp/Zalo diaries, market-stall intercepts — that Zoom-era tooling can’t reach.

Layer three

Routing & Evidence

The expert system that assesses each market’s data richness, routes every study to the right mode, and keeps a ledger where every insight traces to its evidence — field interviews, synthetic pilots, published research, all cited.

The study lifecycle

  1. Intake — decision to be made, markets, segments, timeline.
  2. Market assessment — data-richness scoring per market: language coverage, digital penetration of the segment, prior research corpus.
  3. Routing — expert assigns the mode mix per market and designs the instruments; client sees the routing rationale.
  4. Synthetic pretest — where valid: debug instruments, rank concepts, sharpen hypotheses before spending on fieldwork.
  5. Human fieldwork— the field network collects real responses, met in participants’ channel and language of comfort.
  6. AI synthesis — transcription, translation, theme extraction, cross-market comparison; humans spot-check every load-bearing quote.
  7. Decision-ready readout — recommendations with a full citation trail; every claim clickable back to its evidence.

What the client buys

Which product to build, which market to enter, which pricing survives contact with reality — at the speed of AI where AI is valid, with the confidence of real humans where it isn’t.

Want to build this?

This draft is being shared with friends, operators, and potential partners.

Back to the pitch

Key sources

Bisbee et al. 2024, Political Analysis· Wang, Morgenstern & Dickerson 2025, Nature Machine Intelligence · Gao et al. 2025, PNAS · Cui et al. 2025, Nature Computational Science· Atari et al. 2023, “Which Humans?” (Harvard) · Tao et al. 2024, PNAS Nexus· Sun et al. 2024, “Random Silicon Sampling” · Broska et al. 2025, Sociological Methods & Research · Hullman et al. 2026 · Vaccaro et al. 2024, Nature Human Behaviour · industry positions: NIQ, Kantar, Qualtrics, PyMC Labs.

Working draft — private, for discussion. Not an offer of services.