Reading a problem validation report

Validation is the type that forces an answer, so its report is organised around the hypotheses you wrote. This article covers the parts specific to validation. For the sections every report shares, plus exports, judging how much weight a finding carries, and the thin-run banner, start with Reading a synthesis report.

The five verdicts

Every hypothesis you tested comes back with one of five verdicts. Three of them are results. Two of them are telling you something about the study rather than the hypothesis.

  • Validated. The evidence supports it.
  • Partially validated. It held for some of the audience and not others. Check the per-group read in the appendix before you act, because this often means the hypothesis is true of a segment rather than the market.
  • Not supported. The evidence goes against it. This is a result, and a cheap one.
  • Not applicable to audience. The hypothesis is about something this audience doesn’t do. That is a targeting signal, not a hypothesis signal.
  • Inconclusive. The interviews didn’t produce enough consistent evidence either way. Usually a hypothesis that was too vague to test.

The lead adapts to the result

The headline at the top of What Matters Most is not a fixed template. When the run escalates something, or when the verdicts came back largely inconclusive, the lead changes to say so rather than reporting a tidy result over an untidy study. If the top of your report opens by telling you the study itself is the problem, believe it and fix the hypotheses before reading further.

A map, not a verdict

Under the learnings you’ll find a short note framing the report as a map rather than a ruling. It is worth taking literally. Synthetic participants can tell you a problem statement is coherent, widely recognised, and consistent with how this audience behaves. Confirm with real users before you bet on it. Use validated hypotheses to decide what to build next and what to take to real customers, not as the final word.

What’s in Findings

  • An at-a-glance table. Every problem in one collapsible view, for when you want the shape before the detail. Rows are ranked by severity first, then by how well the evidence behind each one holds up. So the participant count is not the running order: a problem more people mentioned can sit below one with stronger evidence behind it. The Participants column means two different things by row. On a tested hypothesis it is how many participants said the hypothesis is fully true of them, out of those who were asked about it, followed by how many said it is true in part when any did (“2 of 4, 1 in part”). In a card sort, a participant whose answers showed no position counts by where they sorted it. On a problem that surfaced on its own it is how many participants raised it.
  • Tested hypotheses. One row per hypothesis with its verdict and a participant count, each collapsing its own detail and evidence. The count is the thing to read alongside the verdict: Validated at 3 of 12 is a different fact from Validated at 10 of 12. On newer reports the count reads “X of N participants agreed”, adds how many agreed only in part when any did (“X of N participants agreed, Y in part”), and adds how many disagreed when any did. N is the participants who were asked about that hypothesis. That is everyone in a study of up to four hypotheses. In a card-sort study, everyone who sorted the hypothesis counts as asked. A participant whose answers gave no position on it is placed by where they sorted it: very challenging counts as agreed, somewhat challenging as agreed in part, and not challenging or doesn’t apply as not true of them. What a participant said always wins over their sort. A card-sort report built before this rule counted only the participants questioned in depth, so its row can say “of the N participants asked about this”. Open the row’s detail to see who the hypothesis is not true of (listed under Not true of them, which in a card sort includes people placed there by their sort), who agreed only in part, and short quotes of the pushback participants gave when the interview closed by asking what didn’t fit. Pushback aimed at the framing as a whole, not at one hypothesis, is not inside any row. It is its own short list at the bottom of the Tested hypotheses group, titled “Pushback on how the problems were described”, and it appears only when someone gave any. Older reports show “X of N participants” with no “agreed”. That count includes everyone who spoke to the topic, including people who rejected the hypothesis, so don’t read it as agreement. Where a step in Next Steps covers a row, here or in the group below, the row also carries a See Step N link; clicking it opens Next Steps at the step. A row no step covers carries none.
  • Problems that surfaced on their own. Pain raised without being asked, outside your hypothesis set. Each one carries a strength score worked out from the transcripts rather than judged: how many of the participants it was attributed to actually returned to it more than once, and whether they span more than one group. A problem one person mentioned twice reads as weaker evidence than the same problem four people kept coming back to. Every one of them still appears: weakly evidenced problems rank lower, they are never hidden from you. You never see the score itself, on the page or in the download. It only sets the running order. See the question below on why this is often the most valuable section.

In the appendix

Two things sit in the evidence appendix rather than the findings.

  • Sort distribution, present only when the run used a card sort. That happens automatically once a guide carries five or more hypotheses.
  • How each group read the hypotheses. The place to go whenever a verdict came back Partially validated, because it usually resolves into one group agreeing and another not.

Where to go next

Common questions

Usually nothing broke. Inconclusive means the interviews didn't produce enough consistent evidence either way, and the most common cause is a hypothesis that was too vague to test. "Users struggle with onboarding" can't return a clean verdict because every synthetic participant will read it differently. "New users abandon setup because they can't tell which integration to connect first" can. The other common cause is that the hypothesis was outside what this audience does at all, which returns Not applicable to audience rather than Inconclusive. Sharpen the statement and re-run rather than reading Inconclusive as a weak no.

Because it's often where the real finding is. Problems that surfaced on their own collects pain your synthetic participants raised without being asked, outside your hypothesis set. A validation study is by design narrow: you go in with statements and check them. That narrowness is what makes the verdicts trustworthy, and it's also what would make you miss the thing you didn't think to write down. If a problem in this section reached more participants than any of your tested hypotheses, that's your next study.

It appears only when the run used a card sort, which happens automatically once you give the guide five or more hypotheses. Verdicts tell you what held. The sort distribution tells you how your synthetic participants ranked the hypotheses against each other before any deep-diving happened, which is a different signal: a hypothesis can be validated and still have sorted near the bottom for everyone, meaning it's real but not urgent. Read them together when you're deciding sequence rather than just what's true.

More FAQs →

Bring evidence to your next decision.

Start with a free project, or walk through Candor with us first.