Grounding beats the model: why your data matters more than your AI

The model is the part everyone fixates on and the part that matters least. What you ground it in is what decides whether the output is research or improvisation.

Every conversation about synthetic research starts in the same place: which model does it use? Is it on the newest one? Did they switch when the latest version dropped?

It's the wrong question to lead with. The model matters far less than what you feed it. Two tools running the exact same frontier model can produce research-grade output and confident nonsense, and the difference is entirely in the grounding.

Why everyone fixates on the model

It's understandable. The model is the visible part. There's a new headline every few weeks about one model overtaking another, the benchmarks are public, and "we're on the latest model" is easy to say and easy to believe. So the intuition forms: newer model, better synthetic research. Bigger model, more human.

The intuition is wrong, and the research is unusually clear about it.

What the evidence actually says

The largest review of the field looked specifically at model choice across 182 studies and couldn't recommend any model as reliably better for simulating people. Newer and larger did not mean more human-aligned. In some cases an older model produced more human-like cognitive errors than its bigger successor, which is the opposite of the upgrade story. Tweaking model settings like temperature had little effect. And small models fine-tuned on relevant, context-specific data competed with or beat much larger general-purpose ones.

That last point is the one that matters. The lever that moved fidelity wasn't the size of the model. It was grounding it in real, specific data.

Bain reached the same conclusion from the applied side and stated it plainly: the data and context that ground these models matter more than the choice of model. Their guidance to teams building synthetic-customer capabilities put proprietary data first and model selection well down the list.

There's a floor, to be fair. You need a capable modern model, and a weak one will struggle no matter what you feed it. But above that floor, the returns on chasing the newest release are small, and the returns on better grounding are large. Most teams are optimizing the wrong variable.

Why grounding is the part that matters

A model with no grounding has only one place to get its answers: the average of everything it absorbed in training. Ask it about your audience and it hands you the internet's composite opinion, which is fluent, confident, agreeable, and not about your audience at all. That's the flattening problem the skeptics keep pointing at, and a bigger model doesn't fix it. A bigger model just produces a more articulate version of the same average.

Grounding changes where the answers come from. When a persona is built from real evidence about a specific audience (published research, your own customer interviews, survey data, market analysis) and every attribute is tagged with the source it came from, the output stops being the model's generic guess and starts reflecting actual signal. How evidence grounding works walks through the pipeline, and the calibration of personality and bias is part of the same idea: the fidelity lives in the inputs and the structure around the model, not in the model itself.

This is also why "what model is it on?" is a weak way to size up a tool. Every serious tool has access to the same handful of frontier models. None of them own the model. What they own, or don't, is the grounding, the calibration, and the provenance. That's where the real differences are, and it's what the evaluator's framework for synthetic research tools digs into.

What this means for how we built Candor

We treat the model as a swappable part. The underlying model is a setting we can change without touching the architecture, because the architecture is the product, not the model.

The work that makes a Candor persona useful happens around the model: retrieving real evidence before anything is generated, tagging every attribute with its provenance, calibrating personality and cognitive biases from validated distributions, and running a critic pass to catch drift and contradiction. Swap in a better model tomorrow and all of that still does its job, a little better. None of it has to be rebuilt, because none of it was ever riding on which model sat underneath.

That's the deliberate bet. The model is the part that gets better on its own, for free, on someone else's roadmap. The grounding is the part you have to actually build. We put the effort where the difference is.

The takeaway

The model is the easy part. Everyone has the same ones, they improve on their own, and the gap between best and second-best matters far less for synthetic research than the marketing suggests.

The grounding is the hard part, and it's what decides whether you're looking at research or a confident guess in a nice font. If you're evaluating synthetic research, lead with "what is this grounded in, and can you show me the source for a given finding?" instead of "which model is it on?" See how Candor works for our answer. What's the first question you ask when you size up a synthetic research tool?

Common questions

It matters up to a point: you need a capable modern model, and a weak one will struggle no matter what you feed it. Beyond that floor, though, what the model is grounded in matters far more than which model it is. The largest review of the field found no model reliably superior for simulating people, found that newer and larger did not mean more human-aligned, and found that small models fine-tuned on context-specific data competed with much larger general-purpose ones. The lever that moves fidelity is grounding in real, specific evidence about your audience, not the size or recency of the model. Most teams optimize the model and underinvest in the grounding, which is backwards.

Not reliably. The research is unusually clear on this. Across 182 studies, newer and larger models were not consistently more human-aligned, and in some cases an older model produced more human-like cognitive errors than its bigger successor, which is the opposite of the upgrade story. Adjusting model settings like temperature had little observable effect. Chasing the newest release optimizes the most visible variable rather than the one that actually moves output quality. A capable modern model is table stakes; past that, grounding and calibration determine whether the output is usable, not the model's parameter count or release date.

Ask what the output is grounded in (specific, named sources or general AI training data), how personality and bias are calibrated (validated frameworks or unstated assumptions), how provenance is exposed (can you trace a finding to its source), and how consistency is enforced across a study. Every serious tool has access to the same handful of frontier models, so the model is rarely the differentiator. The grounding, calibration, and provenance are where rigorous tools separate from improvisation. If a vendor leads with the model and gets vague on those four, that tells you where they actually invested.

The underlying model is treated as swappable configuration, changeable without touching the architecture. That's deliberate: the architecture is the product, not the model. The work that makes a persona useful happens around the model: retrieving real evidence before generation, tagging every attribute with its provenance, calibrating personality and cognitive biases from validated distributions, and running a critic pass to catch drift and contradiction. When a better model becomes available, it can be adopted without re-architecting, because none of that grounding work was ever riding on which model sat underneath. The model improves on its own; the grounding is the part that has to be built.

Yes, there's a floor. A weak model will struggle to follow instructions, hold a persona, or reason coherently regardless of how well you ground it, so a capable modern model is a real requirement. The point isn't that the model is irrelevant; it's that above that floor, the marginal return on chasing the newest or largest model is small, while the marginal return on better grounding is large. Bigger models do reason better in structured tasks, but they don't fix the core failures of ungrounded synthetic research, like flattening and stereotyping. Grounding does. So the model is necessary but not where the differentiation lives.

Candor is in development.

Be the first to know when it launches.

No spam. Just a note when Candor is ready. Powered by Highline Beta.