Take a study whose human results you already hold. Build personas matched to its population and grounded in real material. Run the same instrument. Compare distributions, subgroup differences and the direction and size of effects, not only means, and normalise against how consistently real people answered on retest. Where the simulation diverges, open the trace and record why. Then publish the result, including the failures. That is a concordance study, and it is the only accuracy figure worth quoting.
The category’s fatal question is “do these predict real people?”, and most answers to it are a correlation from a study the vendor designed, on a population the vendor chose. You can do better in an afternoon. Here is how.
Why validate at all
Because the honest answer to the accuracy question is “it depends”, and the dependencies are exactly the things a vendor number hides: what the persona was built from, which population, which kind of question, whether variance was compared or only averages, and when. The evidence note shows results ranging from 85% of human test-retest for interview-grounded agents to a peer-reviewed failure for prompt-based personas. Both are true. Only a validation on your questions and your buyers tells you which one you have bought.
Even the vendors now say so. One July 2026 buyer guide warns: “Do not rely on universal accuracy numbers or vendor correlation metrics.” ESOMAR, the market research industry body, publishes a buyer checklist titled “5 Topics of Discussion to Help Buyers of Augmented Synthetic Data”. Validation is no longer an unusual demand; it is the expected one.
The concordance method
- 01Choose a study with a known human answer. A published study with public data, or one of your own past studies with the raw results. It must have real respondents, a defined population and an instrument you can re-administer word for word.
- 02Match the personas. Build personas that correspond to the study’s population, grounded in real material about that population, not in a demographic prompt. The Stanford agents were grounded in a two-hour interview each; that grounding is why their result exists.
- 03Run the same instrument. Same questions, same order, same wording. Do not improve the questionnaire; you are testing the personas, not the study.
- 04Compare beyond the mean. Distribution shape, variance, subgroup differences, direction and size of effects, and stability across a repeat run weeks later. The table below is the checklist.
- 05Normalise honestly. Report agreement as a fraction of human test-retest reliability where you have it. Raw correlations flatter and are not comparable.
- 06Read the divergences. Every place the simulation disagrees with the humans, open the trace and record the reason: missing material, a belief that fired, sycophancy, novelty. The reasons are the finding.
- 07Repeat. Run it again after a model update or a material change. Bisbee found the same prompt gave significantly different results three months apart. A validation has a date on it.
What to compare
| Compare | Why | What a pass looks like |
|---|---|---|
| Means | The easy part; a floor, not a proof | Close, but never the only line in the report |
| Distribution shape and variance | Where prompt-based personas fail | Simulated spread within a stated tolerance of human spread |
| Subgroup differences | Segments are the point of research | The same subgroups differ, in the same direction |
| Direction and size of effects | Trends without magnitude mislead | Effects point the same way and are of comparable size |
| Test-retest stability | Results that drift are not results | A repeat weeks later agrees with the first run |
| Novel items | The hardest case | Reported separately, with lower expectations stated |
The two most cited positive studies both used the test-retest yardstick. Stanford’s agents were scored as 85% as accurate as the participants were when matching their own answers two weeks later; the semantic-similarity work reported 90% of human test-retest reliability across 57 surveys. Use the same yardstick and your result is comparable with the literature.
Use the trace where they diverge
A validation that stops at a score tells you how much to trust the instrument. A validation that reads the divergences tells you how to fix it. This is where a grounded persona with a trace is a different tool from a chat log: when the persona disagrees with the human data, its verdict and its “Drawing on” line show whether the material was missing, whether a belief or wound drove the answer, or whether the model simply agreed with the question. Each of those has a different remedy. Missing material you can add. A firing belief you can inspect. Sycophancy you measure, as the pushback note describes.
The score tells you whether to trust it. The divergences tell you why not, and that is the part you can act on.
The discipline is the same one any team running language models in production should already have. Pixelik’s note on evals for business buyers makes the general case; a concordance study is an eval whose golden dataset is a real study.
Publish the failures
A concordance study that only reports where the simulation agreed is marketing. The value, to you and to anyone reading, is in the rows where it did not: which questions, which subgroups, which kinds of novelty, and what the trace said. Publishing them costs a vendor some comfort and buys the only thing the category is short of, which is evidence that can say no.
Perplica’s commitment
No public concordance study of Perplica exists yet, and this site makes no accuracy claim until one does. The next milestone is exactly the method above: three to five published human studies with public results, re-run as Perplica studies against matched, grounded personas, with distributions, subgroups and effect sizes compared against the human data and every divergence read in the trace. The results will be published here in full, failures first. Until then, the label stands: a rehearsal, not a replacement, for talking to real customers.
How do you validate synthetic research?
Run a concordance study: take a study whose human results you already hold, build matched personas, run the same questions, and compare the simulated answers with the human ones on distributions, subgroup differences and direction of effects, not only on averages. Where they diverge, read the trace to see why.
What is a concordance study?
A comparison between simulated and real answers to the same instrument on the same population. The Stanford generative-agent study is one: agents built from interviews were scored on how closely they reproduced the same people’s survey answers, normalised against how consistently those people answered two weeks apart.
Why not compare averages?
Because averages are the easy part. Bisbee and colleagues found synthetic averages close to real ones while variation was lower, regression coefficients differed and results moved with prompt wording and time. A method that only compares means will pass a persona that fails on everything that matters.
What is the right benchmark for accuracy?
Human test-retest reliability: how consistently the same people answer the same questions weeks apart. No simulation can beat it, so reporting accuracy as a fraction of it, as the Stanford and semantic-similarity studies do, is honest. Reporting a raw correlation is not comparable across studies.