Why we don't build digital twins of real people

Candor builds every participant from evidence about a type of customer, never from one real person. That removes most of the consent and privacy questions twins raise, at a cost we explain.

Candor doesn't build digital twins of real people. Every participant in a Candor study is built from evidence about a type of customer, not from any one person's interviews, records or behavior.

That's a deliberate design choice, and it has a cost we'll get to. Two well-funded companies, Simile and Rehearsals, build twins of real individuals, and both describe how the people behind those twins consent. We think they take consent seriously. We just built Candor so that creating a participant never needs one specific person's data or permission. That leaves a lot less to govern.

Two ways to build a synthetic participant

Contrary Research's report AI Polling: How Synthetic Surveys Could Predict Real Behavior, by Claire Burch, sorts the field along one line. It calls the most important difference between these systems whether their simulated populations are based on real individuals.

On one side are twins. The report says Simile maps each agent to one real person, built from structured interviews with that person plus a record of their choices and behavior. Those agents come from about 2.9 million consented responses from more than 400,000 research participants. Simile raised $200 million at a $2 billion valuation in July 2026. CVS Health is a customer and, through CVS Health Ventures, an investor. Simile's own site says every population it builds starts with real people, and that customers can also bring their own data (loyalty data, account histories, telemetry) to train a custom model, on an opt-in basis governed by the customer.

Rehearsals describes its product as twins of real people, where each twin captures one person's decision patterns. Its twins come from an in-house panel. People join through Rehearsals Research, sit a paid 15-minute interview with an AI interviewer, and that interview builds their twin. Rehearsals says companies never receive a participant's video, voice, name or contact details, only modeled responses, anonymized quotes and group-level results. Participants can delete their account, and their twin goes with it. Companies can also use their own customer data to enrich twins.

The report adds that systems like these let a user specify real people: a named public figure, a given customer profile, or a real person picked from a LinkedIn profile.

On the other side are products that build participants who don't exist. That's where Candor sits.

What twinning a real person asks you to govern

Twins have a real advantage, which we'll cover below. They also come with obligations that don't go away just because consent was collected well.

The report names them. Collecting the data needed to twin someone raises ethical concerns for that person. And the output can be turned back on real people: the report warns results could be used to target subgroups, or even specific individuals, "based on responses to surveys that those consumers did not endorse" (Claire Burch, Contrary Research). It compares this to surveillance pricing, where companies use personal data to set a price for you specifically. Zhao and colleagues, cited in the same report, add two worries that apply to any synthetic data: generated answers can be passed off as real, and the way the data was produced has to be open enough for someone else to repeat.

For a research or insights leader, that turns into practical questions. Who agreed to be twinned, to what, and when? If someone withdraws, does their twin disappear from every model that learned from them? What behavioral record sits behind each participant, and who can see it? If you bring your own customer data to build twins, as both companies allow, the consent question now sits partly with you. None of these questions is unanswerable. Rehearsals, for one, publishes answers to several of them, including how a participant deletes their twin. But every one of them is something your legal and procurement teams will want to read.

How Candor builds a participant instead

Candor starts from a description of your audience. The pipeline proposes candidate segments, then gathers evidence about those segments: your own uploaded research first, then web research. It clusters the audience into groups from that evidence.

Each participant is then assembled piece by piece. Their personality scores on the Big Five (the OCEAN model of personality) are drawn from published cross-cultural research, using regional and occupational norms. Their cognitive biases are drawn from ranges set for their group. For U.S. studies, education and household type are drawn in proportion to U.S. Census American Community Survey population shares, so uncommon profiles show up about as often as they do in real life.

The details describing who a participant is, how they behave, what they believe and how they talk each carry a label. Grounded means there's something you can check behind it. Derived means it follows from the evidence without a source stating it outright. Assumed means Candor filled it in and says so. How evidence grounding works walks through the full pipeline.

The result is a participant who resembles a type of customer and matches no real person. Candor doesn't try to predict what a named individual would say, and we don't claim it can.

What that removes from your review

Because no participant is a copy of anyone, a set of problems never comes up for the participants themselves.

  1. No consent chain. There's no panel of real people whose agreement has to be recorded, refreshed and honored when they withdraw. Nobody was twinned, so nobody has to opt out.

  2. No personal behavioral record behind each participant. A participant rests on evidence about a segment and on population-level distributions. There is no file of one person's choices sitting underneath it.

  3. No one gets twinned without knowing it. Candor has no feature for picking a real person, from a LinkedIn profile or anywhere else, and building a participant from them.

  4. No one is targeted from answers they never gave. A participant's answers belong to no real person, so there's no individual behind them to single out.

  5. Less for legal and procurement to review. For teams in healthcare, insurance and financial services, a vendor's data handling gets reviewed alongside the research question, often by compliance. Participants with no real individual behind them give a compliance reviewer a shorter list. If that's your world, see Candor for healthcare CX teams.

There's one exception to that shorter list, and it's the files you bring.

Where real people's data does come in: your uploads

Your own interview transcripts, survey exports, CRM notes, etc. are the strongest evidence Candor can use, and they rank above anything from the web. Your files are read by the AI models that pull out findings. What comes out is written about a type of customer: "nurses put family time above most personal goals," not what one named nurse said. Those findings are what you see in Candor and what participants are built from. Uploaded files are deleted after 30 days, and only the findings stay.

If you want, you can remove names and identifiers before you upload. De-identified research still works as evidence, because Candor looks for patterns across a segment, and a customer's name adds nothing to that.

What twins do better

Twins are more accurate at one specific job.

In the study behind Simile, agents built from two-hour interviews with real people predicted those people's survey answers about nine points better than agents given demographics alone (83% versus 74%, normalized against how consistently people answered themselves two weeks later). If your question is "what will this specific person say?", a twin built from that person will beat anything built without them.

Candor answers a different question: how a type of customer reasons about your problem, and why. That's the job of a synthetic user research study. Evidence about your audience does some of the work an interview does for a twin, and the labels show which details rest on that evidence and which Candor assumed. That makes a Candor participant more than a bare demographic profile, a point synthetic interviews are not synthetic polls covers in more detail.

Where to go next

If you're weighing synthetic research, start with what synthetic research skeptics get right. For how Candor keeps participants from simply agreeing with you, read how Candor guards against people-pleasing participants. And see how Candor works for the full pipeline.

Common questions

A digital twin is an AI agent built to answer the way one specific real person would. It's made from that person's own data, usually a recorded interview plus a record of their choices and behavior. Simile and Rehearsals both build twins of real individuals, and both describe how the people behind them consent. Contrary Research's report on synthetic surveys calls this the most important difference between systems: whether the simulated population is based on real individuals. The other approach, which Candor takes, builds participants who don't exist, from evidence about a type of customer. Only twins need a real person's data and permission to exist.

No. Candor starts from a description of your audience, gathers evidence about the segments in it, and clusters the audience into groups. Each participant is then assembled from that evidence. Their Big Five personality scores are drawn from published cross-cultural research, their cognitive biases from ranges set for their group, and, in U.S. studies, their education and household type from U.S. Census population shares. Participants get invented names. There's no feature for picking a real person, from a LinkedIn profile or anywhere else, and building a participant from them. So there's no consent chain to maintain, and no personal behavioral record behind any participant.

Your interview transcripts, survey exports and CRM notes are the strongest evidence Candor can use, and they rank above anything from the web. AI models read the files and pull out findings. Each finding is written about a type of customer, such as "nurses put family time above most personal goals," not about what one named person said. Those findings are what you see in Candor and what participants are built from. Uploaded files are deleted after 30 days, and only the findings stay. You can also remove names and identifiers before you upload. Candor looks for patterns across a segment, so de-identified research still works as evidence.

At predicting one specific person, yes. In the study behind Simile, agents built from two-hour interviews predicted those people's survey answers about nine points better than agents given demographics alone: 83% versus 74%, normalized against how consistently people answered themselves two weeks later. If your question is what a named person will say, a twin built from that person will beat anything built without them. Candor answers a different question: how a type of customer reasons about your problem, and why. Evidence about your audience does some of the work an interview does for a twin, and each profile detail is labelled to show whether it rests on that evidence.

Bring evidence to your next decision.

Start with a free project, or walk through Candor with us first.