Real users are the ground truth and nothing simulated changes that. Synthetic users are a rehearsal: fast, cheap, available at any hour, useful for narrowing options, generating hypotheses and preparing a study, and, when grounded in real material, accurate enough at the aggregate level to be worth running first. They fail on variance, on subgroups, on genuinely novel propositions and on anything that needs statistical weight. Use them in that order and for those jobs.
The comparison is usually made by whoever is selling one side. Here it is by question type, with the evidence for each cell.
Side by side
| You need to know | Synthetic users | Real users |
|---|---|---|
| The centre of a segment’s opinion | Often good when grounded: 85% of human test-retest in the Stanford agents, 90% in the semantic-similarity work | The reference; slower |
| The spread inside the segment | Weak: less variation than real surveys in Bisbee et al. | This is the point of fieldwork |
| Reaction to something genuinely new | Unreliable: trends yes, magnitude and variability no | Required |
| Why someone reacted as they did | Good if there is a trace; nothing if there is only a chat log | Good, if they will tell you |
| The follow-up they would refuse | You can ask it | You cannot |
| A decision-grade number | No | With a proper sample |
| Speed and cost | Minutes, low | Weeks, high |
Where synthetic earns its place
Nielsen Norman Group, no friend of the category, allows three uses: preparing for studies with real users, treating the output as hypotheses that need testing, and drafting proto-personas and proto journey maps to be refined after real research. Commercial practice has added early screening of a wide option set, message and concept rehearsals, and pricing rehearsals of the kind our playbook describes. The common thread is that each is a step before fieldwork, and each becomes more valuable the more inspectable the persona is: a hypothesis with a trace behind it is a better hypothesis.
The aggregate evidence supports this use. Grounded agents reproduced real survey answers about 85% as well as people reproduce their own answers two weeks later; a semantic-similarity method reached 90% of human test-retest reliability on 57 product surveys. The evidence note lays out the studies and their conditions.
Where only real will do
- Variance and subgroups. Simulated answers cluster toward the centre. Bisbee and colleagues found less variation than in real surveys and regression coefficients that often differed. If the decision turns on how a minority reacts, it turns on fieldwork.
- Novelty. A model trained on the past reproduces how people reacted to things like this before. A genuinely new proposition is, by definition, not like anything before. NN/g’s review found synthetic users capture trends but not the magnitude of effects.
- Surprise. Real participants have priorities and dislike things. Synthetic users, in NN/g’s phrase, seem to care about everything. The reaction you did not predict is the one a simulation is least able to produce.
- Decision-grade numbers. Willingness to pay, share of preference, conversion: these need samples and sampling frames. A simulation is not a sample of anyone.
- Empathy. NN/g’s verdict stands: synthetic users “cannot replace the depth and empathy gained from studying and speaking with real people.”
A simulation is a rehearsal for a conversation with real people. It is a good rehearsal exactly to the degree that it is honest about being one.
The two-phase stack
The pattern that is settling across the category is sequential. Phase one, synthetic: take a wide set of concepts, messages or price points, run them past grounded personas, read the traces, and cut the set to the survivors with written hypotheses about why. Phase two, human: put the survivors to real participants with the sample the decision deserves. The synthetic phase pays for itself if it removes one round of fieldwork on options that were never going to survive, and it improves the human phase by arriving with sharper questions.
Even the platforms selling synthetic panels describe it this way in their customer stories: test early with synthetic respondents, validate final decisions with people. The order is not a compromise. It is the method.
Time and cost, honestly
Vendors quote large savings. Qualtrics, launching its synthetic panels in March 2025, claimed research time cut “from weeks to minutes” and costs reduced “by as much as 70%”. Take the direction and discount the magnitude: the saving is real when synthetic work replaces a round of fieldwork on options that would have failed, and imaginary when it replaces the round that would have told you the truth. The cost that vendors rarely mention is the material. A grounded persona needs your transcripts and documents, and gathering them is work. It is also the work that makes the output worth anything.
What Perplica is for
Perplica is built for phase one done properly: personas grounded in your material, a reasoning trace on every reply so a hypothesis arrives with its why, focus groups where personas react to each other, and a report where every finding cites the turn behind it, with its own limitations printed at the end. It is not built to replace phase two, and every exported report ends by saying its personas are AI simulations, not real customers.
Should I use synthetic users or real users?
Both, in order. Synthetic users to narrow a wide option set, generate hypotheses and prepare a study; real users to decide. The published evidence supports simulated answers at the aggregate level and warns against them for variance, subgroups and novel questions.
What can synthetic users do that real users cannot?
Answer at any hour, at any scale, and take the follow-up question a real participant would refuse. They can also be rerun with one trait changed to see how fragile a finding is. None of that makes them right; it makes them useful before fieldwork.
What can real users do that synthetic users cannot?
Surprise you. Real participants have priorities, dislike things, contradict the moderator and each other, and react to genuinely new propositions in ways a model trained on the past cannot. Nielsen Norman Group’s verdict is that synthetic users cannot replace the depth and empathy of speaking with real people.
Does Perplica replace market research?
No. It is a rehearsal, not a replacement, for talking to real customers. The product and the terms say so, and every exported report ends by stating that personas are AI simulations. It is built to make the rehearsal auditable: grounded personas, a reasoning trace on every reply, and findings that cite their turns.