Fifteen questions to ask a synthetic research vendor, with the answers to listen for

The questions that separate a research instrument from a chat with a system prompt: source material, grounding, traces, refusal, validation and data handling.

R-0923 Sept 2026Updated 28 Sept 20268 min readBy Ikou Soufyane · Founder
The short answer

Ask what the personas are built from, whether every answer shows its reasoning and cites what it drew on, whether personas ever decline or push back and whether anyone measures it, what validation exists against real human data on questions like yours, and where your material goes. The answers to listen for are specific and inspectable. The answers to walk away from are fluent, general and end in a percentage with no study attached.

Every vendor in this category demos well, because the models underneath are fluent. Fluency is not the thing you are buying. Here are fifteen questions that get underneath it, grouped by what they test, with the answer that should satisfy you and the one that should not. The same discipline Pixelik applies to AI vendors in general, narrowed to this category.

Source material

  1. 01What is a persona built from? Listen for: your transcripts, documents and briefs, attached to the persona. Walk away from: “a detailed prompt” or “our proprietary data” with no way to see it. The grounded versus prompted note explains why this question decides everything downstream.
  2. 02Can I see the material a persona holds, and how well it is supposed to know it? Listen for: a list of attached sources, each with a declared level of fluency. Walk away from: a black box.
  3. 03Show me a persona declining a question its material does not cover. Listen for: an in-character “I would not know that”, or a flagged assumption. Walk away from: a confident answer to a question the material could not possibly answer.

Reasoning and receipts

  1. 01Does every answer show why the persona answered that way? Listen for: a stored verdict and trace per reply, inspectable for past turns. Walk away from: “you can ask it to explain”, which produces a second fluent answer, not a record of the first.
  2. 02Does a reply show what it drew on? Listen for: per-turn source attribution, the mechanism the PersonaCite paper calls response-level attribution. Walk away from: sources listed once, at persona level, or not at all.
  3. 03In the report, does each finding link to the exact reply behind it? Listen for: a link from every evidence item to its turn, and a rule that evidence without a turn is dropped. Walk away from: a summary with quotes you cannot trace.

A vendor who cannot show you a trace is asking you to trust a fluent paragraph. That is the one thing this category should never ask.

Pushback

  1. 01Do personas ever disagree with me, and what makes them? Listen for: a mechanism, such as a trust state that starts neutral and is earned, with grounding that gives the persona something specific to disagree from. Walk away from: “you can prompt them to be critical”, which produces theatre.
  2. 02Do you measure the pushback rate? Listen for: a measurement, compared across trust levels, that they can walk you through. Walk away from: an assurance with nothing measured behind it. The sycophancy research says untuned assistants agree by default; a vendor who has not measured its personas against that default has not checked.

Validation

  1. 01What have you validated against real human data, and on what? Listen for: a named study, the population, the human benchmark, the question type, whether distributions and subgroups were compared, and the date. Walk away from: one percentage. Even vendors’ own buyer guides now warn against universal accuracy numbers.
  2. 02Can I run a concordance study on my own past data before I buy? Listen for: yes, and here is how. Walk away from: a case study instead. The validation method describes what to run.
  3. 03How do results change when the underlying model changes? Listen for: metering of model calls and a re-validation habit. Walk away from: a blank look. Peer-reviewed work found the same prompt gave significantly different results three months apart.

Data handling

  1. 01Where does my material go, and is it used to train anyone’s models? Listen for: named subprocessors, a contractual no-training commitment, and the policy in writing. Walk away from: “it is secure”.
  2. 02Can I export everything and delete everything myself? Listen for: a self-serve export of all research content and a self-serve deletion that removes the account and everything under it. Walk away from: a support ticket.

The report

  1. 01Does the report state its own limitations? Listen for: a limitations block the system writes at compile time, naming sample size, the simulated nature of personas and the scenarios not tested. Walk away from: a report that reads as certain.
  2. 02What does the report say personas are? Listen for: a plain statement that they are AI simulations, not real participants, and a term of use forbidding presenting simulated output as real responses. Walk away from: silence, which is a liability waiting for a stakeholder to discover it.

ESOMAR’s buyer checklist, “5 Topics of Discussion to Help Buyers of Augmented Synthetic Data”, covers adjacent ground from the industry body’s side; read it alongside this list. The concordance method turns question nine into something you can run yourself.

How Perplica answers its own list

Built from your material, attached per persona with a declared fluency; three grounding modes, and strict personas decline beyond their material in character. A verdict and trace stored on every reply, inspectable for any past turn; “Drawing on” per reply; every report finding cites its turn or is dropped. Trust starts at 50 and is earned; the pushback rate is measured across trust levels, and the figure will be published with the first concordance study. Validation: no public concordance study yet, and no accuracy number claimed until there is one; the first study is the next milestone and you can run one on your own data. Subprocessors named in the privacy policy, no training on your traffic, self-serve export and deletion. Every exported report ends with its limitations and the line that personas are AI simulations, not real customers. The product page shows each of these from live captures, and the questions above are how to check that any of it is true.

Common questions
What should I ask a synthetic research vendor?

Five things: what the personas are built from, whether every answer shows its reasoning and cites its sources, whether personas push back and whether that is measured, what validation exists against real human data on your kind of question, and how your material is handled by the models underneath.

What is the single most revealing question?

“Show me a persona declining a question its material does not cover.” A prompted persona cannot do it; it will invent a plausible answer. A grounded persona will say, in character, that it does not know. The demo answers the grounding question in ten seconds.

Should a vendor give an accuracy percentage?

Only with the study behind it: population, human benchmark, question type, whether distributions were compared and when it was run. A percentage on its own is a marketing number. Both ESOMAR’s buyer checklist for augmented synthetic data and vendors’ own 2026 guides now say the same.

Is my research material used to train the vendor’s models?

Ask, and get it in writing. In Perplica, conversation content and persona profiles are sent to Anthropic and OpenAI to generate replies and embeddings, and Perplica does not permit that traffic to be used to train their models. The privacy policy states it.