Skip to content
All articles

CONTRIBUTED ARTICLE

What a synthetic survey replication score is measured against

By Alexander Doudkin, CEO, Minds

Several synthetic research companies have recently published on the same question: how do you know a simulation is right? One benchmarks against how much a survey disagrees with itself when fielded twice. One argues that matching answer shares is easy to fake. One scores its own confidence.

They are circling the question in our title: a synthetic research score means nothing until you know what it is measured against. But the harder question is not whether one number looks high. It is whether the same synthetic population can reproduce distributions, preserve individual differences, explain why people answered as they did, and hold up against data you already trust.

That is the problem Minds is built around. Minds is not just a synthetic panel provider. We build, ground and validate the Audience and every individual Mind inside it, using public statistics, research context and, where available, customer data, with answers citing the sources they draw on. The Audience then persists, reusable across quantitative and qualitative work. So validation runs on three levels: the Audience's distributions against real surveys, each Mind's fidelity to a real person, and the whole Audience against research you have already fielded. Here are the seven comparisons we hold it to.

What the number measures

The usual score is aggregate approximation: compare the share of synthetic and real respondents choosing each answer, average the gaps, subtract from 100. If real people split 25/35/30/10 and the synthetic panel 20/40/25/15, every answer is off by 5, so the score is 95%.

It tells you what share said what. It cannot tell you why, and the why is usually what research is paid for. It is also an error measure dressed as accuracy, so it needs a reference point.

Against predicting nothing

At the Audience level, start with the floor. Give every answer an equal share, with no model and no data, and that flat guess already scores far above zero: over 95% on a near-even yes/no question. At the top, even the same survey fielded twice never matches itself exactly. A score is only readable between those two points.

We measured the floor for four public replications, which test the methods behind Minds PRISM, the engine beneath every Mind, in a controlled evaluation setup:

Study Predicting nothing Minds Share of the gap closed
PISA 2022, five countries 79.46% 93.80% 70%
CES 2024, US election 86.14% 95.28% 66%
GSS 2024, US social 83.71% 92.47% 54%
ANES 2024, US election 82.95% 91.67% 51%

Minds closed half to 70% of the distance from a flat guess to a perfect score. That share, not the headline percentage, is what we want to be judged on, because it cannot be inflated by choosing easy questions.

Against the real person

The second level is the individual. A distribution can match without anyone in it being right, so we tested individuals: 22 people, 66 of their later interview answers withheld, each approach answering as that person would, and automated judges, blind to the source, comparing every answer with what the person actually said. We ran the held-out interview test against frontier models:

Approach Fidelity (0 to 100) Picked as closest to the real answer
Minds, grounded in the person's earlier answers 67.66 72.7%
Generic GPT-5.4 42.17 16.7%
Same model, prompted generically 35.24 3.0%
Generic Claude Sonnet 5 37.61 0.0%

Minds was picked eleven times as often as the chatbots' average and more than four times as often as GPT-5.4. Its answers were deeper too: 67 for specificity against 27 to 31, and 32 for genericness against 70 to 76, where lower is better. The same model prompted generically won 3% of the time, so the gain comes from grounding, not from the model. A chatbot gives the answer anyone might give; a Mind gives the answer that person would.

Against a room of experts

Individual fidelity has to add up to range. On idea diversity, a panel of individually profiled Minds scored 7.90 out of 10, against 8.90 for human experts and 3.14 for a uniform AI baseline: close to 90% of the experts' range (study).

Against the public's ranking

Often only the order matters. For Super Bowl LX, an Audience of 100 Minds on the current platform, grounded in Census, American Community Survey, Gallup and General Social Survey data, rated all 54 ads. Against the USA TODAY Ad Meter's roughly 190,000 panelists, the rank correlation was 0.85, and in 83% of head-to-head pairs Minds picked the same winner. Ratings ran 3.87% of the scale generous, so: use the ranking, calibrate the level.

Against a flattering baseline

Asking for probabilities beats forcing a single pick by double digits: 12.5 to 19.7 points in our studies. But in all four, the forced pick scores below the flat guess (PISA 77.71% against 79.46%, CES 73.07% against 86.14%, GSS 74.26% against 83.71%, ANES 77.22% against 82.95%). So quote the gain against the floor, not the forced pick. It is also why each Mind gives a probability for every answer. This concerns rebuilding distributions from single picks, not choice-based methods such as MaxDiff.

Against the other markets

For claims about differences between countries, the fair baseline is the other countries. In PISA, averaging the other four scores 93.42%; Minds scores 93.80%. That is strong fit across five countries, not yet proof of what separates them. We suggest applying this test to every multi-market claim, ours included.

Against the question your team is asking

Which comparison matters depends on who asks. Market research needs distributions. Product and UX research needs the why: reasons, hesitations, follow-ups, reactions to a prototype. Marketing needs all of it, plus A/B tests and a read on the live website.

Distribution replication is only one layer of that workflow. In Minds, an Audience is not generated for one study and thrown away: it is built, grounded and validated once, then reused. The same Minds answer the survey, take one-to-one interviews and follow-up questions on any answer, and react to ads, video, images, mock-ups, live websites and, where enabled, Figma designs, through A/B comparisons, usability and concept tests and methods such as MaxDiff. One validated population, reused across research methods, rather than a new sample for every question.

Seven questions to ask any vendor

  • What does a flat guess score, and how much of the gap does the result close?

  • Where is the individual-level evidence, and can I ask a respondent why?

  • Does the panel give a real range of views?

  • How often does it pick the same winner as real people?

  • What was the comparison actually against?

  • What happens if you predict a group from the others?

  • Can the same validated population give me the distribution, the reasons and the A/B read, or is it rebuilt for every study?

Our validation checklist covers each in more depth.

Against your own data

The third level is the one only you can run. Every Audience in Minds has a Validation tab: upload a survey you have already fielded, questionnaire and results as PDF, Excel, CSV, Word or text, and Minds asks the same Audience the same questions and scores how close it comes, with a 95% range. It can also check against published surveys that match the Audience, or score a study you already ran at no extra cost, and new Audiences can be validated automatically once their Minds finish training. Where a decision needs recruited-human observation, physical or sensory testing, regulated evidence, representative population estimates or final high-stakes validation, those supplement the workflow.

The ultimate benchmark isn't ours. It's yours. Much of this market still sells through a contact form and a scoped pilot. Minds is self-serve: build an Audience, run it against research you have already fielded, inspect where it matches and where it does not, then ask those same Minds why they answered the way they did.

Full studies, methods and limitations: getminds.ai/research. Analysis scripts and per-question numbers behind every figure are available on request.

Alexander Doudkin is CEO of Minds (getminds.ai), which builds, grounds and validates synthetic Audiences and the individual Minds inside them for commercial research.