The 85 to 95% accuracy claim, unpacked: seven numbers, seven different measurements

Every vendor quotes a number in the same band. Each one measures something different on a different task. Where each figure comes from, and what it can carry.

R-1102 Oct 20269 min readBy Ikou Soufyane · Founder
The short answer

The band is real and it is not one number. 85 is the fraction of human test-retest accuracy that interview-grounded agents reached on survey items. 88 and 85 to 92 are two vendors’ self-reported similarity scores. 90 is a median Spearman correlation across 53 survey questions in one consultancy replication, and separately the fraction of test-retest reliability a semantic-similarity method reached. 95 is a client’s description of how similar two sets of conclusions felt. Different metrics, different tasks, different systems. Quoted as one range for the category, the band means nothing.

Ask five vendors how accurate synthetic respondents are and you will hear a number between 85 and 95. Ask what it measures and the answers stop agreeing. This note traces each figure to its source and says what each can and cannot carry.

The band

The range appears in buyer guides, sales decks and vendor FAQs as if it were a property of the technology. It is closer to a coincidence: several unrelated measurements, on unrelated tasks, happened to land in the same ten points. One vendor buyer guide now says so directly: “Vendor benchmark scores cannot serve as universal accuracy metrics because test conditions, baseline datasets, prompt architectures, and evaluation criteria differ across vendors.”

Where each number comes from

Each figure, traced
FigureSourceWhat was measuredWhat it can carry
85%Stanford, Park and colleagues, 2024Agents built from two-hour interviews with 1,052 people reproduced their General Social Survey answers 85% as accurately as the people reproduced their own answers two weeks laterThe best-evidenced figure in the band, and it requires grounding in a real person’s interview
85 to 92%Synthetic Users, vendor site“Synthetic-organic parity in independent comparison studies. Measured across thematic overlap, depth & qualitative alignment”A qualitative overlap score defined by the vendor; the studies are not public
88%Evidenza, vendor site“88% accuracy in 100+ validations”, described as an average similarity score between synthetic and traditional researchSelf-reported; metric and samples not published
0.90EY with Aaru, October 2025EY recreated its 2025 Global Wealth Research Report in one day; “evaluated across 53 single-select questions, achieving a median Spearman correlation of 0.90”A real replication on aggregate rankings; population and questions chosen by the client
90%Maier and colleagues, 2025A semantic-similarity elicitation method reached 90% of human test-retest reliability on 57 consumer surveysA method result on purchase-intent items, not a product claim
95%EY’s CMO, quoted by Evidenza“The conclusions were 95% the same, the correlation was very strong, and in many cases the numbers were nearly identical”A client’s characterisation of one brand survey comparison, not a measurement
0.95Dillion and colleagues, 2023, via MeasuringUCorrelation between synthetic and human moral judgments, with “many points with large differences” between individual ratingsHigh aggregate correlation alongside large individual divergence, in one sentence

Read down the third column and the word “accuracy” dissolves. One row is a fraction of human test-retest reliability, which is the honest yardstick because no simulation can beat how consistently people agree with themselves. One is a rank correlation across survey items. Two are vendor-defined similarity scores with unpublished studies behind them. One is a quotation. The evidence note covers the peer-reviewed results in this table in more depth.

Why they are not the same number

  • Different yardsticks. A Spearman correlation of 0.90 on 53 aggregate rankings and 85% of test-retest accuracy on individual survey items are not comparable, and neither is comparable to “thematic overlap”.
  • Different grounding. The Stanford agents each carried a real person’s two-hour interview. EY’s simulation drew on “demographic, behavioral and sentiment data sourced from national censuses, financial institutions and social media”. A vendor demo may carry a paragraph.
  • Different questions. Survey rankings and stated preferences are the easy end. Magnitudes, novel concepts, subgroups and anything longitudinal are where the record turns.
  • Different authors. Three figures come from research papers, one from a consultancy’s own replication, two from vendors’ marketing pages, one from a customer quote.

A number without its yardstick, its grounding and its question is a decoration. The band is seven decorations in a row.

What the critics measured instead

The critical literature does not dispute that aggregates can match. It measures the things a decision actually rests on. Bisbee and colleagues found synthetic survey data with “less variation in responses than in the real surveys” and regression coefficients that often differed significantly. A 2026 cross-domain benchmark on the General Social Survey and the World Values Survey found that at the individual level “no LLM beats even the strongest baseline”, that models treat demographics as far more predictive of attitudes than they are among real people, and that on a segment-targeting task they “inflate between-segment gaps two to fourfold” and would send a team to the wrong segment in half of the US cases. Tigre and Souto, the same year, listed the failures that aggregate metrics conceal: “variance compression, coefficient sign-flips” and subgroup errors of 10 to 30 percentage points, and proposed diagnostics for deciding when to trust, correct or abandon synthetic panel data.

MeasuringU’s review of the experimental record, published in April 2026, reaches the buyer’s conclusion: “until there is strong evidence of consistently good matching with human data, it seems premature to rely on research with synthetic users for critical decision-making.”

How to read the next one

  • Ask for the yardstick. Fraction of human test-retest is comparable across studies. A raw correlation or a “similarity score” is not.
  • Ask for the grounding. What were the personas built from, and would yours be built from the same kind of material?
  • Ask for the spread. Means are the easy part. Variance, subgroups and effect sizes are where the 2024 to 2026 record is negative.
  • Ask who ran it. A vendor’s own page, a client’s quote, a consultancy’s replication and a peer-reviewed paper carry different weight, and the band mixes all four.
  • Run one yourself. The concordance method takes an afternoon and answers the only version of the question that matters: on your questions, with your material.

What Perplica quotes

Nothing from the band. No public concordance study of Perplica exists, so no accuracy figure appears anywhere on this site, and the label reads as it should: a rehearsal, not a replacement, for talking to real customers. The first concordance study is the next milestone and will be published in full, with the yardstick, the grounding and the questions stated, because after this note we could hardly do otherwise. The vendor questions turn the checklist above into things to ask in the room.

Common questions
Are synthetic research tools 85 to 95% accurate?

Some systems have produced figures in that band on specific tasks: 85% of human test-retest accuracy for interview-grounded agents on survey items, a median Spearman correlation of 0.90 across 53 survey questions in one consultancy replication, and self-reported similarity scores of 85 to 92% and 88% from two vendors. They measure different things and none transfers to your question without a test.

Where does the 90% figure come from?

Two places. EY recreated its 2025 Global Wealth Research Report with Aaru in one day and reported a median Spearman correlation of 0.90 across 53 single-select questions. Separately, a semantic-similarity method reached 90% of human test-retest reliability on 57 consumer surveys. Both are aggregate-level results.

Where does 95% come from?

Mostly from a client characterisation: EY’s Chief Marketing Officer said of an Evidenza brand-survey comparison that “the conclusions were 95% the same”. A 2023 study also found a 0.95 correlation between synthetic and human moral judgments while individual ratings diverged widely. Neither is a measured accuracy of the category.

What do the critical studies find?

That aggregates can match while everything a decision depends on does not: less variation than real people, regression coefficients that flip sign, subgroup errors of 10 to 30 percentage points, individual-level predictions no better than simple baselines, and segment gaps inflated two to fourfold. Those are the 2024 to 2026 results from Bisbee, Chen, Tigre and their colleagues.