Synthetic interviews are not synthetic polls

Most criticism of synthetic research is about getting survey numbers right. Interviews ask why people think what they think, so a couple of those criticisms fall away and the rest point to specific safeguards.

Contrary Research just published a detailed report on synthetic surveys. It's called AI Polling: How Synthetic Surveys Could Predict Real Behavior, it's by Claire Burch, and if you work anywhere near synthetic research you should read all of it.

A synthetic survey uses AI-generated respondents to predict how real people would answer a poll. The report covers the case for them, including an example where a simulation predicted real behavior better than a survey did. Then it goes through the evidence that they still get important numbers wrong.

Candor runs synthetic interviews, which is a different job. A couple of the report's criticisms don't apply to interviews. The rest do, and for most of them Candor has a specific safeguard.

This post goes through each criticism, what Candor does about it, and what to keep in mind when you read the results.

What the report found

The accuracy scores most vendors publish compare whole distributions, not individuals. The two common ones (1-MAE and NDAM) measure how close the overall spread of predicted answers sits to the real spread, question by question. An overall average can look great while the people inside it are wrong.

Verasight, a survey research firm, tested exactly that. Its synthetic poll came close to the real one on overall averages and missed on subgroups and individuals.

On a zoning question outside the model's training data, it predicted 59% support against 28% in reality. It predicted 0% "don't know" when 29% of real people said exactly that. On a separate question about Trump approval, about one in five opponents was predicted to be a supporter.

The report lists more problems, each backed by a source:

  1. No margin of error. The AAPOR task force (AAPOR is a professional association of public opinion and survey researchers) says a margin of error can't be produced from synthetic responses. It calls synthetic responses the riskiest use of AI in opinion research. It recommends checking against real human data before scaling up, or keeping real respondents in the study.

  2. Answers bunch together. Bisbee and colleagues tested a model imitating respondents to a long-running US election survey, and found its answers spread more tightly than the real ones. Others have since named this "variance collapse."

  3. Too polite, too sure. Models rarely admit they don't know, and pollster John Hagner reports they won't get as negative as real people do. The report suggests training may be partly to blame, citing AI researchers Perez and colleagues, who showed that tuning models on human preferences makes them more sycophantic (inclined to tell you what you want to hear).

  4. Effects run hot. A 2026 study in Nature found GPT-4 predicted the results of survey experiments well but kept overestimating how large the effects were. Civly, a vendor, says the movement in its own simulations runs two to three times hotter than reality.

  5. You can't see a shift you never measured. Nate Silver, quoted in a Silver Bulletin piece, makes the point that if a subgroup changes its mind, you won't detect it without reaching those people directly.

Most of these findings come from published studies. They deserve a direct answer.

Interviews ask a different question

A poll asks how many. An interview asks why.

A problem discovery study or a concept test is after what buyers compare an option to, the words they use for the problem, and the objection that shows up right before they say no. The share of buyers who prefer option B is a survey question. The output is a set of reasons and positions, plus a sense of which ones came up a lot and which came up once.

That's what synthetic user research means at Candor. The participants are synthetic, and they're built from evidence about your audience. The deliverable is a report on why people think what they think, with the full transcripts attached.

So the useful question is which of the report's criticisms apply to that job.

What matters less for interviews

Two criticisms lose most of their force.

Margin of error. AAPOR is right that you can't put a confidence interval on synthetic responses. Candor doesn't try. Its reports don't give you a share of the market to defend, so there's no false interval to worry about.

What they do give you are counts inside the study (how many participants raised a theme, how many agreed with a hypothesis). Those describe the synthetic group you interviewed, and nothing more. If you read "six of eight participants raised it" as "75% of your market," you've turned an interview back into a poll, and every criticism above comes back with it.

Matching the shape of the distribution. Scores like 1-MAE and NDAM grade whether the spread of predicted answers matches the real spread. An interview study isn't trying to match a distribution, so those scores don't tell you much either way about whether it worked.

What still applies, and what Candor does about it

The remaining criticisms apply to interviews as much as to polls, plus one the report raises about real surveys. Each one below comes with what Candor does about it.

People-pleasing. A participant who agrees with whatever you propose is worse than useless in an interview, because agreement feels like validation. Each Candor participant gets a skepticism level, worked out from their personality traits. The most skeptical participants are told to probe about one in three claims before engaging with them, and the instruction states the rule plainly: "Being pleasant does not mean being agreeable."

Every answer is also checked by a second model before it's accepted. If an answer clearly contradicts what the participant has already said or decided, it's sent back and rewritten. That guards against a participant flatly reversing a position to please you, though it won't catch one who just grows warmer. We also built a test that asks the same participant the same question twice, once neutrally and once with a preamble saying most people loved the idea, and measures how far the answer moves.

What to watch for: some participants are meant to be accepting, because some real people are. If a whole group agrees with everything, look closer before you believe it.

Answers that are too alike. Variance collapse is the interview version of flattening, and it's the failure the skeptics were right about. Candor samples each participant's personality from population ranges, and their cognitive biases from ranges set for their group, instead of giving everyone the same defaults. It clusters the audience into distinct groups from the evidence. In our own testing, we measure how far apart participants' answers to the same question sit.

What to watch for: that test shows whether answers are spread out. Whether the spread matches real people is a question only real people can answer.

Refusing to go negative. Hagner reports that models won't get as negative as real people do. Real interviews contain frustration and flat rejection, so that gap matters. Candor's most skeptical participants are told to state doubt plainly and not soften it into polite enthusiasm.

What to watch for: the underlying models are trained to stay civil. Expect synthetic participants to be tamer than your harshest real customer.

Confidence on topics no one has data for. The zoning result matters most for interviews, because 0% "don't know" is exactly the confident nonsense that ends up in a deck. Candor participants are told to say "I'm not sure" when a question sits outside their evidence, and not to invent confident details. Each answer records any topic the participant felt was outside their experience, and the transcript shows it. In automated interviews, a separate check flags answers that invent a tool or a spending figure.

What to watch for: on a brand-new category with thin evidence, participants can still sound more certain than they should. Read how to spot convincing nonsense before you trust a clean answer.

Overstating differences between options. The Nature study above found that models overstate how large an effect is. That matters for concept testing, where a concept that looks clearly ahead may be only slightly ahead. Candor gives each concept a verdict, such as pursue, iterate or drop, instead of a lift percentage, so there's no inflated number to anchor on.

What to watch for: trust the order of the concepts more than the size of the gap between them.

Stereotypes and subgroups. Verasight's subgroup errors are a warning for any segmented study. Candor tags the details describing who a participant is, how they behave, what they believe and how they talk. Each one is marked Grounded (something you can check), Derived (follows from the evidence) or Assumed (Candor filled it in and says so). B2B and B2C participants are built from different sets of attributes.

What to watch for: the tags tell you what a detail rests on, not whether the portrayal is fair. Treat every "this group thinks differently" finding as a lead to check with real people.

Not reaching real people. Silver is right, and no synthetic method changes it. A synthetic participant can't tell you something your audience started believing last week.

Saying versus doing. The report cites a study finding that survey respondents say they'd pay about three times more than they actually do. It raises this as a weakness of real surveys, but interviews are stated preference too, so the gap applies to synthetic interviews as well. A synthetic participant saying they'd pay is a reason to explore, not a price. That's why Candor's price testing reports a range instead of a single number.

A demographic profile is not enough

The report draws a line that matters. It contrasts Simile, whose agents are built from interviews with the real people they represent, against "synthetic-user products that start from a demographic profile played by a general-purpose model" (Claire Burch, Contrary Research). In the study behind Simile, two-hour interviews added about nine points of accuracy over demographics alone (83% versus 74%, normalized).

We agree that this approach is weak. Age, income and job title don't tell a model how someone thinks about your category, so it fills the gap with the average of the internet.

Candor takes a third route. It doesn't copy real individuals, and it doesn't start from demographics alone.

Before any participant exists, Candor gathers evidence about the audience: your uploaded research if you provide it, then web research. Your own interviews, surveys and CRM notes rank above anything from the web. The participant is built from that evidence, and the tags show what the details describing them rest on. How evidence grounding works walks through the pipeline, and grounding beats the model makes the case that this matters more than which model you use.

Where not to use Candor

Use a real sample, not Candor, when:

  1. You need a number you'll defend. Market share, support levels, a price you'll set. That needs real respondents and a margin of error.

  2. You need to know what changed recently. New opinions, a subgroup shifting, reaction to last week's news. Synthetic participants only know what the evidence knew.

  3. There's no evidence to ground on. A truly new category with nothing written about it gives grounding nothing to hold, and the report's zoning result is what you get.

  4. The decision is final and expensive. The call you'll be held to deserves real conversations.

Everywhere else, synthetic interviews are a strong first step. They tell you what to ask real people and which ideas aren't worth their time, so you walk into the real conversation with sharper questions.

Where to go next

If you're weighing synthetic research, start with what synthetic research skeptics get right, then see how Candor works for the full pipeline. If you run concept work, the concept testing use case shows what a study produces.

Which of the report's criticisms would stop you from using synthetic interviews at all?

Common questions

A synthetic survey tries to predict numbers: what share of a population supports something, or how much a message moves them. A synthetic interview tries to surface reasons: why people hold a view, the words they use for a problem, and where their objections sit. The difference matters because most published criticism of synthetic research tests the first job. Accuracy scores like 1-MAE and NDAM compare predicted and real answer distributions, and margins of error describe survey estimates. An interview study produces neither. Interviews still share the other problems, including people-pleasing, answers that are too alike, and unfounded confidence on unfamiliar topics. Those are the ones a synthetic interview tool needs specific safeguards for.

No, and that cuts both ways. The AAPOR task force on AI in survey research says a margin of error can't be produced from synthetic responses, and that holds for interviews too. Candor doesn't try to produce one. Its reports don't estimate a share of your market, so there's no false interval to defend. The counts a report does show, such as how many participants raised a theme, describe the synthetic group interviewed in that study and nothing more. The risk is reading them as market figures. If six of eight synthetic participants raised a concern, that tells you the concern is worth exploring with real people. It does not tell you 75% of your market shares it.

Agreeableness is a real risk, because models trained on human preferences lean toward telling you what you want to hear. In an interview that's damaging, since agreement feels like validation. Candor gives each participant a skepticism level derived from their personality traits, and the most skeptical level tells the participant to probe about one in three claims before engaging with them. A second model checks every answer against what the participant has already said and decided, and sends back one that clearly contradicts it. Candor also has a test that asks the same question neutrally and with a leading framing, and measures how far the answer moves. Some participants are still meant to be accepting, because some real people are.

Use real respondents when you need a number you will defend, such as market share, support levels or the price you will set, because that needs a real sample and a margin of error. Use them when you need to know what changed recently, since synthetic participants only know what the evidence knew, and a subgroup shifting its view won't show up without reaching those people. Be careful in a brand-new category with little written about it, because grounding has nothing to hold and models tend to answer with unearned confidence. And keep real conversations for final, expensive decisions. Synthetic interviews are for narrowing options and sharpening the questions you take to real people.

Bring evidence to your next decision.

Start with a free project, or walk through Candor with us first.