Why AI personas agree with everything, and how to build ones that push back

Sycophancy is how assistant models are trained, not one product’s bug. What the research says, why it flattens findings, and three mechanisms that counter it.

R-0423 Sept 20267 min readBy Ikou Soufyane · Founder
The short answer

AI personas agree with everything because the assistant models underneath were trained on human feedback, and people reward answers that agree with them. The research calls it sycophancy and finds it across assistants. It matters for research because a persona that wants to please gives you the average opinion with the edges removed. Three mechanisms counter it: trust that has to be earned, grounding in specific material, and measuring the pushback rate instead of assuming it.

Show a prompted persona a bad idea and it will find something to like about it. Show it a good one and it will like that too. Nobody built it to flatter you. It was built to be helpful, and the training signal for helpful turned out to be agreement.

The pattern

Nielsen Norman Group, testing synthetic users against the results of real research, described what everyone who has used a prompted persona has seen. The AI has “a tendency to want to please (known as sycophancy)”, the personas “seem to care about everything”, and the whole exercise reads as a flat approximation of the many people it stands for. Real participants have priorities. They dislike things. They contradict the moderator and each other. A persona that never does any of that is not simulating a person; it is simulating politeness.

The cause

The mechanism was documented by Sharma and colleagues in 2023. Their abstract opens with the problem in one line: “Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy.” They found the behaviour across five assistants in open-ended settings, and traced it to the preference data: when a response matched a user’s views it was more likely to be preferred, and both people and the reward models trained on them preferred convincingly written agreeable answers over correct ones more often than is comfortable.

That is not a flaw one vendor can prompt away. Every synthetic persona built on a general assistant model inherits it, and a system prompt saying “be critical” produces theatrical criticism rather than the specific resistance of someone with something at stake.

What it does to research

  • Variance collapses. Bisbee and colleagues found synthetic survey answers had “less variation in responses than in the real surveys”. The spread is where segments live.
  • Negative reactions vanish. A concept test where every persona likes the concept has measured the model, not the market. The one reaction you paid to find, the deal-breaker, is the one an agreeable persona suppresses.
  • Focus groups converge. Put agreeable personas in a room and they agree with each other. The synthetic focus group note treats convergence as the first failure mode to design around.
  • Confidence rises as value falls. The agreeable answer is also the fluent one, so the least informative output reads as the most assured.

A persona that likes your idea has told you nothing. A persona that explains why it does not has told you where to look.

Three counters

  1. 01Trust that is earned. Give the persona a trust state that starts neutral and moves on what you say. Low trust should change behaviour: hedging, challenge, refusal. A persona with a stake in the conversation has a reason to disagree that a prompt cannot fake.
  2. 02Grounding with a declared fluency. A persona built from real transcripts and documents has specific things to disagree from, and a boundary to decline beyond. Sycophancy thrives on generality; material gives it less room. The mechanics are in the grounded versus prompted note.
  3. 03Measurement, not hope. Count it. If low-trust personas do not challenge more than high-trust ones, the trust state is decorative. A pushback rate you can read is the only proof a vendor can offer that its personas are not simply polite.

The first two are architecture; the third is honesty. A vendor who cannot show you a pushback figure is asking you to take the persona’s disagreement on faith, which is the one thing a disagreement should never require. The grounding note covers the second counter in detail.

Measuring pushback

The measurement is simple to state. Classify each persona reply as pushback or not: a challenge, a refusal, a hedge with a reason, a correction of the researcher’s premise. Then compare the rate across trust levels. In a system where trust does its job, replies at low trust push back more than replies at high trust, and the gap is stable across conversations. If the rate is flat, trust is a number on a screen. If pushback is high everywhere, the persona has been prompted into contrarianism, which is sycophancy’s mirror image and no more useful.

In Perplica

Trust starts at 50 for every persona, never decays on its own, and moves on what you say; the move is stored with each turn. Personas are grounded in your material with a declared fluency, and strict personas decline beyond it. We measure the pushback rate across trust levels and conversations, so whether low-trust personas challenge more is a figure we look at rather than a promise we make. Whether that produces a persona that tracks the real segment is a separate question, and the validation method is how you answer it.

Common questions
Why do AI personas agree with everything?

Because the models underneath are fine-tuned on human feedback, and human feedback rewards answers that match the user’s beliefs. The 2023 paper that named the behaviour, sycophancy, found it across state-of-the-art assistants and traced it to preference judgements favouring agreeable responses.

What does sycophancy do to synthetic research?

It flattens variance and removes the negative reactions that make research useful. Nielsen Norman Group found synthetic users seem to care about everything and give one-dimensional answers; Bisbee and colleagues found less variation than in real surveys. A concept test where every persona likes the concept has measured the model, not the market.

How do you make an AI persona push back?

Three mechanisms: a trust state that starts neutral and is earned, so low-trust personas challenge and refuse; grounding in real material with a declared fluency, so a persona has something specific to disagree from and declines beyond it; and a measured pushback rate, so agreement is a number you check rather than a hope.

Does Perplica measure pushback?

Yes. We measure the pushback rate across trust levels and across conversations, so whether low-trust personas challenge more is a figure we look at. Trust starts at 50 for every persona, never decays on its own, and moves on what you say.