How group clustering works

Real audiences aren't all one kind of person. They're distinct types who reason differently, and treating them as the same flattens the research. Group clustering surfaces those types from real evidence: 3 to 8 clusters per study, each with calibrated personality and cognitive bias ranges. The synthetic participants you interview are sampled from those ranges, and two from the same group are kept meaningfully different from each other.

This is the technical companion to the earlier piece on how evidence grounding works. Evidence grounding produces the signal pool. Group clustering turns the signal pool into the population structure that drives the rest of the study. For the broader category context, see what is synthetic user research; for the full Candor walk-through, see how Candor works.

What a group is in Candor, and what it isn't

The category vocabulary in synthetic research is fuzzy, and the three terms most often confused are signal, group, and synthetic participant. Each is a distinct object in Candor's pipeline, and they nest in a specific order.

A signal is an atomic piece of extracted evidence. It comes out of the evidence-retrieval stage and represents one structured claim: a behavior, a pain point, an attitude, a constraint, a goal, a belief, a preference, or a decision rule. Every signal carries a provenance tag, and a signal drawn from a web page or an uploaded file shows it: web sources link out and uploads name the file. Most studies produce hundreds of signals.

A group is a cluster of 3 to 8 meaningfully distinct types of person inside the audience, each representing a fundamentally different way of being a member of it. Groups are built in two passes, and only the second one has a name you will see. First the pipeline proposes candidate divisions of the audience straight from your description, and searches for evidence about those specifically. Then it synthesizes the whole signal pool against the dimensions that best separate the audience, and what comes out of that second pass is your groups. The intermediate divisions are working state; the product does not show them to you, and the rest of this essay calls them segments only because that is what the pipeline calls them. Each group has a name, a description, defining dimensions, an OCEAN personality profile expressed as a range per trait, a set of primary cognitive biases with intensity ranges, and a population weight indicating its proportion of the total participant set.

A participant is one respondent sampled from a group. Each participant inherits the group's core (defining dimensions, OCEAN range, bias range, memory structure template) and adds independent secondary traits that vary within those ranges. Many participants can share a group, and the system enforces meaningful separation between those siblings so they don't trivially collide on personality or bias values.

The order matters: signals come from evidence, the pipeline's own segments come from clustering those signals at the audience layer, groups come from synthesizing signals across the included segments, and participants come from sampling groups. Each layer is built from the layer below it. Anything a participant says in an interview rests on the group's defining attributes, the signal evidence those attributes rest on, and the source documents and published research that evidence was extracted from. The product shows that chain as labels on the participant's profile fields and source links and file names on the audience's signals, not as a link from each answer back to a document.

What the user controls, and what runs automatically

A useful methodology essay starts by being honest about what the user actually controls. In Candor's group-clustering stage, the user's direct decisions are limited and intentional.

What the user controls:

  • Study setup (audience type, learning goals, industry, region, any uploaded research).
  • Interview-type selection at the interview-guide stage (one of problem discovery, problem validation, concept testing, or price testing).

What runs automatically inside the group-clustering pipeline:

  • Signal summarization (turning hundreds of signals into a structured implication pool).
  • Dimension selection (picking the 5 to 9 dimensions that best separate the included audience).
  • Group generation (synthesizing the groups themselves).
  • Critic validation (checking the groups for realism, distinctness, evidence alignment, and segment coverage).
  • Respondent sampling (drawing the specific OCEAN and bias values for each participant from the group's ranges, with sibling-distance enforcement).
  • Memory generation (building each participant's six memory structures from the signal evidence).

The user can inspect the dimensions and the candidate trait pool on the audience-review screen if they want to understand what the system is working with, but the dimensions that actually drive group clustering are picked by code, not by the user. This is deliberate. Letting users hand-pick clustering dimensions in a low-volume research context usually produces groups that confirm what the team already believes, which defeats the point of running the research at all. Automating the dimension selection from evidence-anchored variance makes the output less convenient to manipulate and more useful as a research instrument.

The pipeline at a glance

The group-clustering pipeline is one phase of participant generation, which runs straight after audience generation in the same background job. The clustering stages run in this order:

  1. Signal summarization.
  2. Dimension selection.
  3. Group generation.
  4. Critic validation.
  5. Respondent sampling and memory generation.

The first four stages produce the groups themselves. The fifth stage produces the individual participants by drawing from the group-level ranges. The rest of this piece walks through each.

Stage 1: From signals to clusters

The signal pool that comes out of evidence retrieval is rich but unstructured for clustering purposes. Hundreds of behavioral, attitudinal, and contextual signals, each tagged with provenance and, where one was recorded, its source, do not yet describe a clustered audience. They describe a population's behavior, attitudes, and decisions in fragments.

The first thing the participant-generation pipeline does is summarize those signals into a structured implication pool that's usable for clustering. The summary preserves the diversity of perspectives in the underlying signals (so the group stage doesn't collapse the audience into a single "average" view) and translates raw signal language into the categories that downstream clustering can reason about: which behaviors recur across the audience, which attitudes split the audience into recognizably different camps, which constraints differentiate one segment of buyers from another.

This summarization step is also where the pipeline starts to surface the shape of the audience that the segments alone don't show. Segments describe demographic and behavioral clusters; signals capture the underlying reasoning patterns. Two segments that look similar demographically can be quite different in how they reason about a decision, and the signal-summarization stage exposes those reasoning differences as inputs for the clustering work.

Stage 2: Choosing the dimensions that separate the audience

Before generating groups, Candor picks the dimensions that will actually separate one group from another. This is the trickiest part of the methodology, because the answer to "which dimensions matter" depends entirely on the audience.

A consumer audience for a wellness product might cluster cleanly on values, lifestyle, and risk tolerance. A B2B audience for a developer tool might cluster on engineering maturity, build-vs-buy disposition, and team size. A regulated CX audience might cluster on technology comfort, plan tenure, and prior experience with care navigation. Picking the wrong dimensions produces groups that look distinct on paper but reason about decisions identically.

The dimension-selection step runs deterministically inside the participant-generation pipeline. It pulls from the candidate trait pool that audience generation produced (the user can browse this pool on the audience-review screen, but doesn't pick from it) and selects 5 to 9 dimensions that:

  • Show high variance across the included segments (so the groups will actually look different).
  • Are grounded in the signal evidence (so the dimensions aren't generic stereotypes).
  • Distinguish on reasoning patterns, not just demographic surface (so the groups capture different ways of thinking, not different ways of looking).
  • Span the relevant category mix for the audience type (B2B emphasizes role, decision environment, organizational context, and constraints; B2C emphasizes lifestyle, identity, and aspirational self).

The output is a set of 5 to 9 selected dimensions, each with a short rationale for why it was picked. The group-generation stage uses these as the scaffolding for clustering.

The user doesn't choose these dimensions. The selection is automatic, deterministic, and grounded in evidence variance. The reason is the one already mentioned: in low-volume research contexts, letting users pick clustering dimensions tends to produce confirmation bias, not insight.

Stage 3: Generating the groups

With the selected dimensions in hand, the pipeline synthesizes the groups themselves. This is the most LLM-heavy step in the clustering pipeline, and it produces structured output rather than free text.

For each group, the generation step produces:

  • A name: a short, human-readable label (for example, "The Cautious Evaluator," "The Tooling Skeptic," "The First-Time Buyer"). Names are intended to be memorable and to capture the group's core stance, not to be flippant.
  • A description: two or three sentences summarizing the group's defining stance, current behaviors, and decision pattern.
  • Defining dimensions: the 3 to 5 dimensions where this group is most distinct from the other groups in the set. Each dimension carries a value and a distinctiveness indicator (how much this dimension separates this group from peers).
  • A full trait profile: the values for all of the selected dimensions on this group, not just the defining ones. This is the complete behavioral and attitudinal shape.
  • An OCEAN profile: a range for each of the five Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). Ranges, not points. Each participant sampled from this group gets specific OCEAN values drawn from inside these ranges.
  • Primary biases: the cognitive biases that are most relevant to this group's decision-making, each with an intensity range (for example, status-quo bias 70 to 90). Candor's bias library covers twenty-one biases across shared, B2B-specific, and B2C-specific categories. Intensity ranges, not points, for the same reason as OCEAN.
  • Signal sources: which signals from the evidence pool informed this group's construction. The audit trail at the group level.
  • A population weight: the proportion of the total participant set that this group should represent. Population weights sum to 1.0 across the group set and drive how many participants get allocated to each group downstream.

How many groups get generated per study depends on two things: the number of included segments and whether the audience is niche. Niche audiences cap at 4 groups, because forcing 7 distinct types out of a thin population is a rigor problem, not a feature. A narrow audience of 1 or 2 segments produces 5 groups, a mid-range study with 3 to 4 produces 6, and a broad audience of 5 or more produces 7.

The groups are generated as a coherent set, not one at a time. The system reasons about all of them together so they end up meaningfully different from each other rather than minor variations on the same core type.

Stage 4: Critic validation

The generated groups go through an automatic critic agent before they're persisted. The critic checks the group set against four criteria:

  • Realism: does each group describe a plausible person given the signal evidence, or does any group combine traits in ways that don't track to documented behavior?
  • Distinctness: are the groups meaningfully different from each other on at least three dimensions, or are some groups too close to others to be useful as separate research targets?
  • Evidence alignment: do the defining dimensions on each group ground in the signal pool, or has the generation step invented traits that have no support in the underlying evidence?
  • Segment coverage: do the groups collectively represent every part of the audience the pipeline found evidence for, or has the set drifted toward one part at the expense of others?

The critic is structurally similar to the critic in the audience-generation pipeline. It runs as a separate validation step rather than as part of generation, and when it identifies issues, the system addresses them before continuing rather than passing weak output downstream. The user sees only the final, critic-passing group set, which is one reason group output tends to read as more coherent than a single-pass LLM persona generator would produce.

Stage 5: Sampling participants from group ranges

Once the group set passes critic validation, the pipeline generates individual participants by sampling from the group-level ranges. This is where the move from cluster-level structure to instance-level synthetic respondents happens.

The total is four participants for every group in the set, capped at 24. Those slots are then shared out in proportion to each group's population weight, so a group carrying 30% of the population gets more participants than one carrying 10%, while every group gets at least one (the minimum-coverage rule).

Each participant's sampling process:

  • OCEAN values are drawn from the group's OCEAN ranges. Two participants in the same group get distinct OCEAN profiles, never the same point in personality space. The system enforces meaningful separation between siblings on OCEAN so two in the same group don't accidentally end up nearly identical.
  • Bias intensities are drawn from the group's bias ranges. Same approach: participants in the same group get different bias intensity values, so a study with four participants in one group produces four meaningfully different decision-reasoning profiles.
  • Secondary traits (demographic, B2B firmographic or B2C lifestyle fields, behavioral specifics) are sampled with secondary-trait variation, so even siblings have distinct surface identities.
  • Memory structures (six memory types per participant: identity, behavioral, belief, language, decision, and conversation memory) are generated from the signal evidence plus that participant's sampled traits. The conversation memory starts empty; the others are populated at generation time and updated as interviews accumulate.

The sibling-distance enforcement is one of the methodology details that separates group clustering from naive replication. Without it, sampling four participants from a group with overlapping OCEAN ranges can produce four statistical near-twins, which collapses the research breadth that having several participants per group is supposed to provide. The system prevents this at sampling time by checking proposed OCEAN profiles against already-sampled siblings and resampling if they collide.

B2B versus B2C: structural separation, not lip service

Candor models B2B and B2C as fundamentally different research domains throughout the group-clustering stage, not just at the cosmetic level. The differences show up in four specific places:

Dimension emphasis. B2B group dimensions emphasize organizational context, decision environment, role and responsibility, and external constraints (vendor lock-in, security review processes, procurement cycles). B2C group dimensions emphasize lifestyle, identity, aspirational self, and lived-experience constraints. The dimension selection stage uses different category weights depending on whether the study is B2B or B2C.

Cognitive bias baselines. B2B and B2C audiences have different default bias intensity baselines. Status-quo bias, sunk-cost fallacy, career-risk aversion, and authority bias run higher at baseline in B2B because the decision-making environment rewards these biases (changing tools is risky for the buyer's career, established vendors carry implicit authority). Optimism bias, present bias, and scarcity or FOMO effects run higher at baseline in B2C because consumer decisions are made under different incentive structures.

The private-versus-public-stance distinction (B2B specifically). B2B participants maintain two layers of belief: their actual private view of a topic, and their likely committee-public stance on the same topic. These can differ. A B2B participant might privately think a competing vendor has the better product but publicly advocate for the incumbent because supporting the existing choice protects their relationship with the procurement team. This duality is baked into B2B group memory structures.

Decision-rule shape. B2B decision rules tend to include explicit committee dynamics (who needs to approve, in what order, with what evidence). B2C decision rules tend to include impulse and identity considerations. The system models these differently rather than treating them as variations of the same underlying decision framework.

The result is that B2B groups and B2C groups don't just have different content. They have different structural shapes, which produces interview behavior that reflects how each audience actually reasons rather than producing B2B interviews that feel like consumer interviews with formal language pasted on top.

What group clustering gives you, and what it still can't do

The cumulative effect of the five stages is a population structure that's grounded in evidence, audited at the cluster level, and varied from one participant to the next. Specifically:

Meaningful breadth in your research. Four participants in the same group aren't statistical near-twins; they're four different reasoning profiles within the same broad type. A five-group study with 20 participants gives you 20 meaningfully different interview perspectives, not five perspectives repeated four times each.

Provenance from cluster to instance. When a participant says something distinctive in an interview, you can see what the participant rests on (their profile labels each field) and open the signals behind their group, each showing its source where it came from a web page or an uploaded file. The audit trail crosses every layer.

Calibrated psychology, not labels. Personality and bias enter the model as ranges that get sampled, not as binary tags on the participants. A participant with "high anchoring bias" actually reasons differently from a sibling with "moderate anchoring bias," because the bias intensity is a real continuous parameter, not a flag.

Honest population structure. The group set represents the audience your description and evidence describe, in proportions that reflect that audience, with weighting visible to the researcher rather than hidden in averages.

What group clustering can't do:

It can't capture genuine outliers below the population-weight threshold. A 5-group model collapses minor variations into the nearest group. If your audience contains a small but strategically important subgroup that doesn't justify its own group slot, that subgroup is folded into the cluster nearest to it and loses some fidelity. The fix is to include that subgroup as a separate audience segment at Phase 1, not to expect clustering to recover it.

It can't update the group core from interview feedback. Participants update their belief memory and decision memory as interviews accumulate within a study. Group-level traits don't update. If the underlying audience drifts mid-study, you'd see it in one participant's belief evolution but not in the group set itself. For longitudinal questions, the right answer is to re-run audience generation with updated evidence rather than expecting in-study drift to propagate up.

It can't model per-domain bias variation. A participant's bias intensity is a single value across all decision domains in the study scope. Real human anchoring bias on price differs from anchoring bias on feature comparisons; the system uses a single intensity for both. This is a fidelity compromise the methodology accepts in exchange for tractable models.

It can't recover from a poorly-defined audience at Phase 1. Group clustering inherits whatever divisions the audience stage produced, and those come from your description. If they are wrong (too broad, too narrow, missing a real subgroup, including a non-existent one), the groups reflect that flaw. The critic agent validates groups against evidence, not against real-world segment structure, so it can't catch this.

It can't model genuinely novel audiences with no evidence support. The clustering pipeline depends on signal evidence to be meaningful. For a brand-new category or a niche professional segment with no public research coverage and no first-party data, the candidate trait pool will be thin and the resulting groups will be flagged as assumed rather than rigorous research output. The honest path in those cases is to either acquire first-party evidence first or treat the synthetic research as hypothesis-grade rather than findings-grade.

Related reading

For the layers above and below group clustering, see the OCEAN model in synthetic participants (how personality calibration is sampled per group) and how participant memory actually works (the six memory types that make a participant behave as the same person across sessions). For the upstream pipeline that feeds group clustering, return to how evidence grounding works. For the broader category context, see what is synthetic user research, the use cases Candor supports, and the comparison hub.

Common questions

Between 3 and 8 by design, and 4 to 7 in practice. The count depends on the number of included audience segments and whether the audience is classified as niche. A niche audience caps at 4 groups, because forcing more clusters than the evidence supports is a rigor problem rather than a feature. A narrow audience of 1 or 2 segments produces 5 groups, a study with 3 to 4 produces 6, and a broad audience of 5 or more produces 7.

A group is a cluster: a type of person, with defining dimensions, OCEAN and bias ranges, and a population weight. A synthetic participant is one individual sampled from that group, with specific OCEAN and bias values drawn from the group's ranges and unique secondary traits. Many participants can share a group, and the system enforces meaningful separation between them, so two in the same group don't trivially collide on personality or reasoning patterns.

No, and that's intentional. The dimension selection step runs deterministically inside the participant-generation pipeline, picking 5 to 9 dimensions from the candidate trait pool based on variance across segments and grounding in the signal evidence. Users can browse the candidate trait pool on the audience-review screen if they want to understand the methodology, but the actual selection isn't a user pick. The reason is that letting users hand-pick clustering dimensions in low-volume research contexts tends to produce groups that confirm what the team already believes, which defeats the point of running the research.

A range allows synthetic participants in the same group to be meaningfully different from each other rather than statistical near-twins. If a group has an extraversion range of 60 to 90, four participants sampled from that group get four different extraversion values inside that range, and the system enforces minimum separation between them. This produces actual breadth within a group, not just four copies of the same personality profile. Cognitive bias intensities work the same way: ranges at the group level, distinct sampled values per participant.

Yes, structurally. The dimension selection stage emphasizes different category mixes (organizational context and role for B2B; lifestyle and identity for B2C). The cognitive bias library applies different baseline intensities (status-quo bias and career-risk aversion run higher in B2B; optimism bias and present bias run higher in B2C). B2B synthetic participants maintain a distinction between their private view of a topic and their likely committee-public stance, which doesn't exist in the B2C model. The decision-rule shapes differ accordingly. These aren't cosmetic differences. They produce interview behavior that reflects how each audience actually reasons.

Five honest limitations. It can't capture genuine outliers below the population-weight threshold, which means small strategically-important subgroups get folded into the nearest group. It can't update the group core from interview feedback, so longitudinal drift questions need re-running audience generation rather than waiting for the clusters to evolve. It can't model per-domain bias variation; one bias intensity covers all decision domains. It can't recover from poorly-defined segments at Phase 1, which the user controls. It can't model genuinely novel audiences with no public evidence; in those cases the candidate trait pool is thin and the groups are flagged as assumed rather than research-grade.

Bring evidence to your next decision.

Start with a free project, or walk through Candor with us first.