Skip to content

INDUSTRY GUIDE / 2026 EDITION

Synthetic Market Research 101: State of Industry

An industry guide to synthetic respondents, simulated consumers, synthetic data and digital twins.

What matters now

The conversation around synthetic research has shifted from can a model answer a survey? to how well does it answer one, and how would anybody know? Across the evidence, the same questions recur: where synthetic systems are useful, how their outputs should be assessed and what they cannot replace.

Three things stand out. First, the field produced measurement: total simulated survey error, psychometric audits of model respondents, scales for detecting machine-written answers and preregistered out-of-sample tests of discrete-choice experiments all landed this year. The recurring verdict is that synthetic respondents are plausible but not valid, and that they fail hardest where a study needs a real sampling frame rather than a fluent imitation of one.

Second, "silicon sampling" is a named lineage now, not a novelty, and the interesting failures are structural rather than prompt-level: a well-argued preprint contends that instruction-tuned models cannot faithfully sample from the distributions they are asked to represent, whatever the persona text says. Conditioning ablations benchmarked against real probability samples show where the conditioning effort should go.

Third, the statistics and privacy literature runs years ahead of research practice. Disclosure control, differential-privacy trade-offs and quality checks for low-fidelity synthetic data are settled engineering there, and regulators and statistical agencies frame synthetic data as a privacy tool, not an insight tool. Market-research-facing guidance, such as the silicon-sample guidelines in NIM Marketing Intelligence Review, still reads as provisional.

Digital twins ran on a separate track - industrial, health and urban, heavy on standardisation - with the consumer twin still mostly a promise.

Treat synthetic data as an input that must be evaluated, rather than a substitute for evidence. Decisions should remain clear about the sources, assumptions and validation behind them.

The 2026 evidence map

Theme Published work Directly relevant Open access Videos Public posts
Synthetic respondents (AI-simulated survey participants) 502 272 90% 16 0
Simulated consumers, personas and generative agents 1,167 294 83% 6 0
Synthetic market research, conjoint and choice experiments 1,258 99 73% 5 11
Synthetic data: methods, privacy and official statistics 1,060 252 86% 14 9
Digital twins (consumer, human, urban, industrial) 1,010 754 73% 18 13
Governance, regulation and privacy-enhancing technology 516 45 75% 8 10

Institutions represented

  • Prompt (Canada) — 32
  • Carnegie Mellon University — 27
  • University of Konstanz — 27
  • Cornell University — 25
  • Statnett (Norway) — 24
  • University of Mannheim — 18
  • Innovation Team (China) — 18
  • Huawei Technologies (China) — 17
  • Politecnico di Milano — 16
  • Cobuilder (Norway) — 15
  • The University of Texas at Austin — 13
  • National University of Singapore — 13
  • Office for National Statistics — 13
  • Pro Persona — 13
  • George Mason University — 12
  • Max Planck Institute for Human Development — 12
  • Sichuan University — 12
  • University of Technology Sydney — 12
  • Khon Kaen University — 11
  • Tongji University — 11
  • Centre National de la Recherche Scientifique — 11
  • Tsinghua University — 11
  • Istituto Nazionale di Statistica — 11
  • Stanford University — 10
  • Institut d'Etudes Politiques de Paris — 10

Where the work appears

  • Zenodo (CERN European Organization for Nuclear Research) — 1296
  • arXiv (Cornell University) — 491
  • SSRN Electronic Journal — 170
  • Figshare — 63
  • Lecture notes in computer science — 52
  • Research Square — 38
  • OSF Preprints (OSF Preprints) — 38
  • Mendeley Data — 36
  • Statistical Journal of the IAOS — 35
  • AEA Randomized Controlled Trials — 34
  • Underline Science Inc. — 33
  • ˜The œinternational archives of the photogrammetry, remote sensing and spatial information sciences/International archives of the photogrammetry, remote sensing and spatial information sciences — 24
  • Applied Sciences — 23
  • Preprints.org — 22
  • Scientific Reports — 20
  • Advances in computational intelligence and robotics book series — 17
  • Lecture notes in networks and systems — 17
  • Springer texts in business and economics — 17
  • Sustainability — 16
  • Astronomy and Astrophysics — 15

A. Synthetic respondents (AI-simulated survey participants)

Research on systems that substitute or augment human survey participants with model-generated ones. It includes accuracy audits against real benchmarks, the token-probability/"silicon sampling" lineage and the emerging critique literature.

Selected reading

Watch

In the news

Discussion

Explore the research library

For a deeper reading list, explore the full synthetic respondents research library.

B. Simulated consumers, personas and generative agents

Multi-agent frameworks, persona libraries and social-simulation engines that inform how synthetic respondents are designed and evaluated.

Selected reading

Watch

In the news

Explore the research library

For a deeper reading list, explore the full simulated consumers and personas research library.

C. Synthetic market research, conjoint and choice experiments

Studies that attack research practice directly - LLM-generated conjoint and choice experiments, synthetic survey samples, and the validity questions that follow.

Selected reading

Watch

Public conversation

  • 2026-08-21 · @polpsychangel.bsky.social — Does made up data come close to real data when using LLMs? No. Synthetic data is made up, stop calling it anything more than that. papers.ssrn.com/sol3/papers.... — post
  • 2026-08-12 · @kwcollins.bsky.social — Listening to a podcast about synthetic sample and it’s proponent is bragging about how well it does it back testing. But of course it does well because all of those outcomes are in the training data. — post
  • 2026-08-09 · @arxiv-cs-cl.bsky.social — Alexander Apartsin, Yehudit Aperstein Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies https://arxiv.org/abpost
  • 2026-08-07 · @cscl-bot.bsky.social — Alexander Apartsin, Yehudit Aperstein: Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies https://arxiv.org/apost
  • 2026-08-05 · @hanowell.me — The answer to the question: “can you just make up data and get paid for it” before, yknow, synthetic survey respondents were seriously considered by very serious economists. — post
  • 2026-07-29 · @csspenn.bsky.social — Friday, July 31 Mansfield(210) / Track H / 10:45–12:15 • Can LLMs Approximate Public Opinion? Validating Synthetic Survey Responses Against 40 Waves of Real Survey Data by Elliot Pickens (@elliot-p — post
  • 2026-07-28 · @mattansb.msbstats.info — You'll never guess who came across this post while taking a break from analyzing a giant dataset of synthetic survey data. — post
  • 2026-07-22 · @wendynorris.bsky.social — The dangers of LLMs and synthetic data in survey work. Westwood, S. J. (2025). The potential existential threat of large language models to online survey research. Proceedings of the National Academ — post
  • 2026-07-07 · @ihi-synthia.bsky.social — Your voice matters. Help shape the future of synthetic data in healthcare by taking part in our European survey. Share to help us reach more citizens across Europe. 🗣️ Read more here: bit.ly/4c7if0o — post
  • 2026-06-25 · @adamdrummond.bsky.social — There's a scenario where market research shifts en masse to synthetic data with the odd human survey to keep the models fresh. Then the cost of research with real people gets much more expensive or di — post
  • 2026-06-25 · @sophieehill.bsky.social — I just don't get it... If a "synthetic survey" generated an accurate result at T1, it would either be luck or good aggregation of existing survey data. Why would we expect it to give an accurate res — post

In the news

Discussion

Explore the research library

For a deeper reading list, explore the full synthetic market research research library.

D. Synthetic data: methods, privacy and official statistics

The statistics and privacy discipline that market research mostly ignores: disclosure control, evaluation metrics, synthetic populations and benchmarking.

Selected reading

Watch

Public conversation

  • 2026-08-31 · @feed.thedigitalspeaker.com.ap.brid.gy — In an era of AI hallucinations, synthetic media, and misinformation, validation determines whether AI creates value or catastrophe. Organizations without systematic validation are one bad output away — post
  • 2026-08-19 · @bigearthdata.ai — Sharing more, protecting more: Three lessons from the Safe Data Technologies project ->Brookings / More on "Privacy-enhancing technologies for statistics" at BigEarthData.ai / #Data — post
  • 2026-06-28 · @iam.slys.dev — The gap between real experiments and synthetic data is shrinking. With SynthBH, both contribute to scientific progress, but safeguards stay strong. What new kinds of discoveries might this unlock? st — post
  • 2026-06-16 · @paperposterbot.bsky.social — arXiv📈🤖 An Energy-Driven Framework for Privacy-Aware Synthetic Data Generation By Massoli, Spagnuolo — post
  • 2026-05-14 · @bigearthdata.ai — AI may save lab animals by rescuing small medical studies ->Earth.com / More on "AI synthetic data reduces animal testing" at BigEarthData.ai / #AI — post
  • 2026-04-25 · @data4sci.bsky.social — Designing synthetic datasets for the real world: Mechanism design and reasoning from first principles — Useful perspective on building synthetic datasets by starting fr… https://research.google/blog/dpost
  • 2026-04-23 · @joachimschork.bsky.social — Access to the Statistics Globe Hub April modules is only available to those who join this month: statisticsglobe.com/hub 🔹 Draw Synthetic Data with drawdata in Python 🔹 Monte Carlo Simulation 🔹 AI-As — post
  • 2026-04-17 · @joachimschork.bsky.social — K-means clustering is a simple and widely used method for identifying patterns in data. I also use it in a recent Statistics Globe Hub module, where it is combined with synthetic data created using t — post
  • 2026-04-06 · @iam.slys.dev — We're reimagining privacy, access, and scale with synthetic data from AI. But what happens when these generated datasets fail to capture the full messiness of reality? The story is in the gap between — post

In the news

From institutions

Discussion

Explore the research library

For a deeper reading list, explore the full synthetic data research library.

E. Digital twins (consumer, human, urban, industrial)

Digital twins as the industrial-scale cousin of the synthetic respondent - including human, consumer, urban and supply-chain twins, plus standardisation work.

Selected reading

Watch

Public conversation

  • 2026-09-12 · @hiddenscifi.bsky.social — Dan Warner’s sci-fi debut Wither arrives Sept. 29. Regression builds perfect digital twins from every secret and mistake. A year after testing it in an Ozark town, Mick returns to find a ghost town wi — post
  • 2026-09-11 · @arxiv-daily-bot.bsky.social — From State Synchronization to Cognitive Self-Evolution: An Operational Architecture for Cognitive Digital Twins Haoran Gao et al. #arXiv #cs.AI — post
  • 2026-09-11 · @jeromeollier.bsky.social — A high-resolution digital twin of Oeno Atoll (PITCAIRN Islands) through integrated geospatial data - @frontiersin.bsky.social Marine Science www.frontiersin.org/journals/mar... — post
  • 2026-09-11 · @quiltydunn.bsky.social — Follow-up digital twin studies show that off-the-shelf transformers can do the same thing. That is, untrained neural networks, if fed inputs modeled on the inputs received by chicks, can model the dev — post
  • 2026-09-11 · @quiltydunn.bsky.social — We review a recent literature, including work done by Justin's lab, that tightly controls the inputs available to developing animals and to computational models or "digital twins". I think this is s — post
  • 2026-09-11 · @letswinpc.org — PNET patient Burt Rosen experiments with AI to create a digital twin that provides insight into how he might respond to different treatments. https://bit.ly/4xCVlrgpost
  • 2026-09-11 · @noamchompers.bsky.social — The digital twin studies Justin wood et al. have been running are kinda mindblowing, and this is a really serious consideration of their theoretical import — post
  • 2026-09-11 · @shipbot.bsky.social — Inside the Terminal’s Digital Twin: What Real-Time Yard Visibility Looks Like Marine Insight · 11 Sep 2026 — post
  • 2026-09-11 · @crial.bsky.social — La Trasmittanza Assoluta azzera la resistenza d'asse e salda i bilayer fitochimici, il citoscheletro, il Digital Twin e le matrici solide in un unico circuito super-conduttore. Chiedetevi se il vostro — post
  • 2026-09-11 · @aijamesdooley.bsky.social — With a digital twin, entrepreneurs can scale communication, support customers, share expertise, and keep momentum - even when they’re not available every minute of the day. That’s where real leverage — post
  • 2026-09-11 · @aijamesdooley.bsky.social — 📘 Clone Yourself With AI - Why Every Entrepreneur Needs a Digital Twin by AI James Dooley shows how entrepreneurs can use AI to build a digital twin that extends their knowledge, presence, and value w — post
  • 2026-09-11 · @aijamesdooley.bsky.social — Classic Setting, Next‑Gen Mindset - Clone Yourself With AI - Why Every Entrepreneur Needs a Digital Twin Timeless places can spark the most future‑focused business ideas. 1/7 — post
  • 2026-09-11 · @smart-twinvill.bsky.social — Smart TwinVill kicked off today! 14 partners, 8 countries, one mission: digital twins for resilient villages. RAINNO leads Impact Maximisation, dissemination, communication & exploitation, plus manag — post

From institutions

Discussion

Explore the research library

For a deeper reading list, explore the full digital twins research library.

F. Governance, regulation and privacy-enhancing technology

What regulators, statistical offices and standards bodies are saying about generating people and markets in software.

Selected reading

Watch

Public conversation

  • 2026-06-17 · @bigearthdata.ai — Synthetic data generation: challenges and perspectives for gastrointestinal medicine ->Nature / More on "AI synthetic data medical imaging" at BigEarthData.ai / #Data — post
  • 2026-06-02 · @ihi-synthia.bsky.social — 🎧 New SYNTHIA podcast episodes are now streaming on Spotify! Explore conversations on synthetic data, trustworthy AI, privacy, regulation & innovation in healthcare with experts from across the SYNTH — post
  • 2026-04-20 · @joaomatosdigital.bsky.social — The article puts the topic of synthetic users in another level. The level of consent, regulation, data analysis and data science. Synthetic users based on averages will give us avg experiences, not aw — post
  • 2026-03-18 · @synthemaeu.bsky.social — Registration is open for the SYNTHEMA & @erneurobloodnet.bsky.social in haematology webinar series. Join all 4 sessions in May 2026 or register for individual webinars. Synthetic data, federated lea — post
  • 2026-02-16 · @danielwrasmus.bsky.social — Autonomy 🤖 sneaks into workflows—then shows up in the post-mortem. Feb update to 🗓️ The State of AI 2026: rollback governs autonomy, infra is destiny, 🇪🇺 EU AI Act baseline, synthetic data tradeoffs, — post
  • 2026-01-16 · @biorxiv-sysbio.bsky.social — Synthetic Data to Explore Transcriptional Regulation of Differentially Expressed Genes in Ovarian Cancer https://www.biorxiv.org/content/10.64898/2026.01.15.699618v1post
  • 2026-01-16 · @biorxivpreprint.bsky.social — Synthetic Data to Explore Transcriptional Regulation of Differentially Expressed Genes in Ovarian Cancer https://www.biorxiv.org/content/10.64898/2026.01.15.699618v1post
  • 2025-12-17 · @ihi-synthia.bsky.social — SYNTHIA participated in CFE–CM Statistics 2025 in London. Vibeke Binz Vallevik (DNV) contributed to the workshop on synthetic data generation & evaluation, discussing data quality, hallucinations, GDP — post
  • 2025-11-24 · @kai3690.bsky.social — ## Algorithmic Fairness Auditing Through Counterfactual Data Synthesis and Bayesian Network Verification in GDPR Compliance Frameworks Abstract: This paper introduces a novel framework for algori — post
  • 2025-10-30 · @ssrn.bsky.social — "Taxonomizing Synthetic Data" by Cofone et al. classifies synthetic data into transformed, augmented, and simulated types, emphasizing the importance of ground-truth assumptions for legal and policy d — post

In the news

From institutions

Explore the research library

For a deeper reading list, explore the full governance and privacy research library.

How to read this guide

Use this guide to explore the research, examples and debate surrounding synthetic market research. The sections separate synthetic respondents, simulated consumers, choice experiments, synthetic data, digital twins and governance so that readers can follow the questions most relevant to their work.

The evidence shows recurring value in exploration, hypothesis generation and early-stage testing. It also repeatedly points to the need for validation when a decision depends on a real sampling frame, measured behaviour or a representative population.

Research papers, institutional material and public commentary serve different purposes. A published claim about a product or method is a useful starting point for investigation, not proof of its performance in a particular use case. Readers should examine study design, comparison data and the decision context before acting on any individual finding.

Digital twins sit alongside synthetic respondents because the terms increasingly overlap in commercial discussion, even though much of the underlying research comes from industrial, health and urban applications.

CONTINUE EXPLORING

Put the evidence in context.

Explore synthetic market research companies Read institutional guidance Browse all guides