Most research teams using synthetic users are doing it without any rules. Not loose rules. None.
In User Interviews' survey of 150 researchers, 63% said there's no organization-level stance on synthetic users: the call is left to whichever researcher happens to be running the study. This isn't indifference. It's an absence of any shared position, so each person decides for themselves when synthetic is allowed and how far to trust it. Another 15% didn't even know whether a policy existed. So roughly three in four teams are using, or about to use, a method most of them are openly skeptical of, with nothing written down to steer it.
That's not caution. It's a vacuum. And a vacuum gets filled by whoever's most confident in the room, which is usually the person who wants the fast answer, not the careful one.
Why the gap is a problem
The same survey lists what researchers are worried about: the quality of the insights (88%), stakeholders over-trusting AI output (79%), and bias getting amplified across groups that are already underrepresented (77%).
Those are exactly the failures a little governance prevents and no governance invites. If nobody decided in advance how synthetic output gets used, the default is that a clean-looking deck gets treated like panel data, because nobody wrote down that it isn't. The over-trust problem isn't a technology problem. It's a "we never agreed on the rules" problem.
Governance doesn't mean a 40-page policy
It means a short set of defaults, agreed once, that settle two questions before a study runs: when synthetic is appropriate, and how its output is allowed to be used. You can write the first version on a single page. A starter set looks like this.
-
Set the stakes before the study, not after. Decide up front whether the decision is high-stakes, hard to reverse, regulated, or likely to affect a protected or marginalized group. If it is, synthetic can inform the thinking, but it can't be the evidence the decision rests on. The researchers in the survey reached for the same gate: is this a high-risk call like a new product line, and will these insights land on marginalized populations. Answer those two before you generate anything.
-
Require provenance on every finding. "The AI said so" is not a source. Any synthetic finding that makes it into a deck should trace back to the evidence behind it, the same way you'd expect a quote to trace to a transcript. If your tool can't show you where a persona's claim came from, that's the governance gap itself, not a missing nice-to-have. Candor tags every persona attribute with its source for this exact reason, walked through in how evidence grounding works.
-
Say what you grounded it on. Disclose the evidence base in the writeup. "Grounded in our 2025 churn interviews plus public category research" is one sentence. It tells the reader what the findings can carry and what they can't, and it stops the output from looking like it appeared out of nowhere.
-
Anonymize real data before it goes in. If you're grounding synthetic personas on real customer data, strip the identifiers first: names, account IDs, anything that points at a specific person. It's the cheap, responsible default, and it clears most consent and privacy questions before they start. Whatever tool you use, check how it stores, processes, and reuses what you upload. Candor's third-party processor inventory is on the subprocessors page.
-
Backtest before you trust a new use. The first time you point synthetic at a new kind of question, check it against something you already know the human answer to. One researcher in the survey put the alternative bluntly: the only way to check the work was to launch the thing and see if it worked, which is an expensive way to find out you were wrong. A held-out check costs almost nothing by comparison.
-
Label synthetic as synthetic, every time. Every deliverable, every chart, every Slack summary. This is the rule that addresses the 79% over-trust worry head-on. Findings drift loose from their context fast. Six weeks later nobody remembers which numbers came from real people and which came from personas, unless it says so right on the slide.
Who should own this
Research should, and most researchers seem to agree it's theirs to claim.
The striking thing in that survey isn't the 63% gap. It's that the same people naming the gap also see themselves as the ones who should fill it. One senior researcher framed AI bias detection as exactly the kind of work research teams are well-suited to lead. The flip side of "no org-level stance" is open space to set the standard rather than inherit one.
If research doesn't write the rules, someone with less context will. Usually the person who just wants the synthetic number to come back a yes.
Why you should do this now
None of this slows you down. A one-page default set takes an afternoon to write and saves you the same argument on every study after. It also makes synthetic findings easier to defend, not harder, because "here's our evidence base, here's the provenance, here's where we drew the line" is a much stronger position in a stage gate than a confident chart with no audit trail behind it.
The vacuum gets filled either way. Better that you fill it on purpose. For more on why over-trust is the real risk to guard against, see what synthetic research skeptics get right, and for the questions to ask a vendor about data and methodology, the evaluator's framework for synthetic research tools. To see how grounding and provenance work in practice, see how Candor works.
Does your team have written rules for synthetic research yet, or is it still left to whoever's running the study?