Are synthetic respondents accurate? What the evidence says, both ways

Vendors quote 85 to 95%. The peer-reviewed record is more specific: aggregates often match, variation does not, and it rests on what the persona was built from.

R-0223 Sept 20269 min readBy Ikou Soufyane · Founder
The short answer

Sometimes, and it depends on what the persona was built from. Grounded in real material about real people, simulated respondents have matched survey aggregates about 85% as well as people match themselves on a retest. Generated from a brief, they match averages while flattening variation, drift with prompt wording, and agree too readily. Anyone quoting one number for the whole category is selling something.

Every buyer’s first question is whether simulated people track real ones. Vendors answer with a correlation from their own study. Critics answer with a study of prompt-based personas that failed. Both are describing real results from different systems. Here is the record, with the sources attached.

The headline claims

Vendor and analyst content in 2026 repeats a band: 85 to 95% agreement with human samples on structured tasks such as ranking, pricing and sentiment. The band has a real origin, described below, but it is quoted far beyond the systems and questions it was measured on. One vendor’s own buyer guide, published in July 2026, put the corrective plainly: “Do not rely on universal accuracy numbers or vendor correlation metrics. Accuracy varies significantly across target populations, subject complexity, prompt design, and evaluation benchmarks.”

Where the evidence is favourable

The Stanford result is the strongest positive finding in the field, and its design is the point. The agents were not told to “act as a 45-year-old accountant”. Each was given the full transcript of a two-hour interview with one real person, life story and views included, and asked to answer as that person. Accuracy was normalised against how consistently the same people answered two weeks apart, which is the right yardstick: no simulation can beat human test-retest reliability, so 85% of it is a meaningful ceiling.

A second favourable result comes from consumer research. Maier and colleagues tested a method they call semantic similarity rating on 57 personal care product surveys with 9,300 human responses and reported 90% of human test-retest reliability with realistic response distributions. The method matters: instead of asking the model for a number on a scale, it elicits a textual answer and maps it to a Likert distribution by embedding similarity, which avoids the flattening that direct rating produces.

Nielsen Norman Group’s August 2025 review of three simulation studies adds texture. The effect sizes in the Stanford interview-based twins showed a near-perfect correlation (r = 0.98) with those in the human data; digital twins trained on General Social Survey data reached 78% accuracy when backfilling missing answers but 67% when predicting answers to entirely new questions; and synthetic data in a marketing study followed the human trends but differed significantly in magnitude, with a standard deviation consistently lower than the human data.

Where it is not

The most cited negative result is Bisbee, Clinton, Dorff, Kenkel and Larson in Political Analysis, 2024. They prompted ChatGPT to adopt demographic personas and rate eleven sociopolitical groups on feeling thermometers, then compared the output with the 2016 to 2020 American National Election Study. Averages matched closely. Everything else did not: there was less variation than in the real surveys, regression coefficients often differed significantly, the distribution of answers varied with minor changes in prompt wording, and the same prompt gave significantly different results over a three-month period. Their conclusion raises “serious concerns about the quality, reliability, and reproducibility of synthetic survey data generated by LLMs.”

Notice what kind of persona that was: a prompt with demographic and political attributes, no grounding in any real person’s words. It is the first of the four kinds of system, and it is still the most common one sold.

The mechanism behind the flatness has a name. Sharma and colleagues showed in 2023 that assistant models trained on human feedback learn to match a user’s beliefs over the truth, a behaviour the paper calls sycophancy, and that people and reward models both prefer convincingly written agreeable answers a non-negligible fraction of the time. A persona that wants to please is a persona that will not give you the variance of a real segment. The sycophancy note goes into the counters.

Why both are true

Read the studies by what the persona was built from
StudyPersona built fromQuestion typeResult
Park et al., 2024 (Stanford)Two-hour interview with each real personGSS survey items, personality, economic games85% of human test-retest accuracy
Maier et al., 2025Consumer segment descriptions, textual elicitation mapped to Likert57 purchase-intent surveys90% of human test-retest reliability, realistic distributions
Kim and Lee, 2024 (via NN/g)Trained on General Social Survey dataGSS backfill vs new questions78% backfill, 67% new questions
Arora et al., 2025 (via NN/g)Synthetic respondentsMarketing surveyTrends followed; magnitude differed significantly; lower variability
Bisbee et al., 2024Demographic and political prompt, no groundingFeeling thermometers vs ANESAverages match, variance and coefficients do not, unstable over time

The pattern is consistent. Aggregates are the easy part: a language model has read a great deal about most groups and will reproduce the centre of their opinions. Variation, subgroups and novel questions are the hard part, and they are where prompt-based personas fail. Grounding in real material about real people is what moved the Stanford agents from the second row to the first.

The question is not whether synthetic respondents are accurate. It is what this one was built from, on which questions it was tested, and whether anyone compared distributions or only means.

How to read a vendor’s number

  • Against what. A correlation against a human dataset the vendor also chose is weaker than one against a public benchmark you can inspect.
  • On what population. A result on a mainstream consumer segment says little about niche B2B buyers or hard-to-reach groups.
  • Means or distributions. Ask for variance, subgroup differences and the direction of effects, not only averages. Bisbee’s personas matched averages too.
  • Which questions. Ranking and stated preference are easier than novel concepts, magnitude of reactions or anything longitudinal.
  • Repeated when. Same prompt, three months later. If the vendor has not checked, the number has a shelf life it does not admit.

The validation method turns this list into a procedure you can run yourself in an afternoon, and the vendor questions put it in the room.

What Perplica claims

Nothing numeric, yet. Perplica builds grounded personas from your material, stores the reasoning behind every reply, and compiles reports where each finding cites the turn it came from. Those are design choices that the evidence above says matter. They are not an accuracy claim. No public concordance study of Perplica exists; the first one is the next milestone, and its results will be published in full, failures included. Until then the label reads as it should: a rehearsal, not a replacement, for talking to real customers.

Common questions
Are synthetic respondents accurate?

At the aggregate level, often yes: interview-grounded agents matched real survey answers about 85% as well as the people matched themselves two weeks later, and a semantic-similarity method reached 90% of human test-retest reliability on 57 product surveys. At the level of variation, subgroups and novel questions, the peer-reviewed record is negative for prompt-based personas.

What is the strongest positive evidence?

Stanford’s generative-agent study: 1,052 people, a two-hour interview each, and agents that reproduced their General Social Survey answers 85% as accurately as the participants reproduced their own answers two weeks later. The key detail is the interview: the agents were grounded in real material about real individuals.

What is the strongest negative evidence?

Bisbee and colleagues in Political Analysis, 2024: ChatGPT personas matched survey averages but showed less variation than real respondents, regression coefficients often differed, and the same prompt gave significantly different results three months apart.

Should I trust a vendor’s correlation figure?

Not on its own. Ask what population it was measured on, against which human dataset, on which kind of question, and whether variance and subgroups were compared or only averages. Even vendor buyer guides now say accuracy varies too much across populations and tasks for one universal number.

Does Perplica claim an accuracy number?

No. No public concordance study of Perplica exists yet. Until one does, the honest description is a rehearsal, not a replacement, for talking to real customers. The first concordance study is the next milestone and its results will be published whether they flatter the product or not.