Every conversation about synthetic research starts in the same place: which model does it use? Is it on the newest one? Did they switch when the latest version dropped?
It's the wrong question to lead with. The model matters far less than what you feed it. Two tools running the exact same frontier model can produce research-grade output and confident nonsense, and the difference is entirely in the grounding.
Why everyone fixates on the model
It's understandable. The model is the visible part. There's a new headline every few weeks about one model overtaking another, the benchmarks are public, and "we're on the latest model" is easy to say and easy to believe. So the intuition forms: newer model, better synthetic research. Bigger model, more human.
The intuition is wrong, and the research is unusually clear about it.
What the evidence actually says
The largest review of the field looked specifically at model choice across 182 studies and couldn't recommend any model as reliably better for simulating people. Newer and larger did not mean more human-aligned. In some cases an older model produced more human-like cognitive errors than its bigger successor, which is the opposite of the upgrade story. Tweaking model settings like temperature had little effect. And small models fine-tuned on relevant, context-specific data competed with or beat much larger general-purpose ones.
That last point is the one that matters. The lever that moved fidelity wasn't the size of the model. It was grounding it in real, specific data.
Bain reached the same conclusion from the applied side and stated it plainly: the data and context that ground these models matter more than the choice of model. Their guidance to teams building synthetic-customer capabilities put proprietary data first and model selection well down the list.
There's a floor, to be fair. You need a capable modern model, and a weak one will struggle no matter what you feed it. But above that floor, the returns on chasing the newest release are small, and the returns on better grounding are large. Most teams are optimizing the wrong variable.
Why grounding is the part that matters
A model with no grounding has only one place to get its answers: the average of everything it absorbed in training. Ask it about your audience and it hands you the internet's composite opinion, which is fluent, confident, agreeable, and not about your audience at all. That's the flattening problem the skeptics keep pointing at, and a bigger model doesn't fix it. A bigger model just produces a more articulate version of the same average.
Grounding changes where the answers come from. When a persona is built from real evidence about a specific audience (published research, your own customer interviews, survey data, market analysis) and every attribute is tagged with the source it came from, the output stops being the model's generic guess and starts reflecting actual signal. How evidence grounding works walks through the pipeline, and the calibration of personality and bias is part of the same idea: the fidelity lives in the inputs and the structure around the model, not in the model itself.
This is also why "what model is it on?" is a weak way to size up a tool. Every serious tool has access to the same handful of frontier models. None of them own the model. What they own, or don't, is the grounding, the calibration, and the provenance. That's where the real differences are, and it's what the evaluator's framework for synthetic research tools digs into.
What this means for how we built Candor
We treat the model as a swappable part. The underlying model is a setting we can change without touching the architecture, because the architecture is the product, not the model.
The work that makes a Candor persona useful happens around the model: retrieving real evidence before anything is generated, tagging every attribute with its provenance, calibrating personality and cognitive biases from validated distributions, and running a critic pass to catch drift and contradiction. Swap in a better model tomorrow and all of that still does its job, a little better. None of it has to be rebuilt, because none of it was ever riding on which model sat underneath.
That's the deliberate bet. The model is the part that gets better on its own, for free, on someone else's roadmap. The grounding is the part you have to actually build. We put the effort where the difference is.
The takeaway
The model is the easy part. Everyone has the same ones, they improve on their own, and the gap between best and second-best matters far less for synthetic research than the marketing suggests.
The grounding is the hard part, and it's what decides whether you're looking at research or a confident guess in a nice font. If you're evaluating synthetic research, lead with "what is this grounded in, and can you show me the source for a given finding?" instead of "which model is it on?" See how Candor works for our answer. What's the first question you ask when you size up a synthetic research tool?