Reading a concept testing report

A concept test report is built to drive a build-or-pivot call, so it leads with a verdict per concept and works backwards to the evidence. This article covers the parts specific to concept testing. For the sections every report shares, plus exports, judging how much weight a finding carries, and the thin-run banner, start with Reading a synthesis report.

Read the alerts first

Up to four callouts can appear near the top of What Matters Most, just under the headline. Each one is telling you that a result further down needs reading differently.

  • Some concepts were misread. Your participants didn’t understand what the concept was. Any verdict on a misread concept is a verdict on a misunderstanding, so fix the description before you conclude anything about the idea. On a name, value proposition, or messaging test the heading says names, value propositions, or messages instead.
  • Some names set the wrong expectation. Participants understood the concept, but its name promised something the description didn’t deliver, so the fix is to rename, not to rewrite the description. On a name test it means some participants didn’t think a name fit the product: it promised something the product doesn’t do, suggested the wrong kind of product, or gave no clue what it does, and the box gives no rename advice. For each flagged name the box gives three counts: participants who said the name doesn’t fit, participants who said it fits, and participants who gave no view on whether it fits. On a standard test, that last count includes participants who never explored the concept in depth. Older reports show one count instead, such as “11 of 23 said the name doesn’t fit.”
  • Some value propositions weren’t believed or Some value propositions were misread on a value proposition test, and Some messages weren’t believed or Some messages didn’t come across on a messaging test. When the box holds both problems, its heading says the options didn’t land. For each flagged option it counts two things apart: participants who misread it, and participants who understood it but didn’t believe it. An option is listed only when those two together outnumber the participants who found it persuasive. The counts stand alone, with no example of what a participant said. They need different fixes. A misread needs a clearer rewrite. A doubt needs proof, or a claim you can back up. Doubting a claim still counts as understanding it.
  • This may be an audience-fit problem. The concept was understood, but it didn’t land with the audience you tested while a different group showed real pull. This one is framed as an opportunity rather than a warning, and it seeds the first next step. Treat it as a lead to validate, not a conclusion.

The first three link straight to the concept they refer to, further down in Findings. The audience-fit one doesn’t link. When the report can tell which concept the group pulled toward, the box names that concept, and a Drop on it means don’t advance it with the audience you tested. If any of these boxes is present, resolve it before you act on a Drop.

The decision table

Also in What Matters Most: one row per concept, ordered best to worst, each with its verdict, how many participants named it their favourite (for example 18 of 24). The favourite count appears there and under the concept’s Intent score, nowhere else. When the concept with the most favourites is not the one to pursue, the headline box says so in a sentence of its own, and the next sentence gives the reason it is not the one to pursue now. A concept with more favourites never shows a lower Intent grade than one with fewer. When the grade was raised to match the favourite count, the count sentence under it is the whole explanation. The next step for each concept sits in Next Steps, and the concept’s own row further down in Findings links to it. If you read nothing else, read this. It is the whole study compressed into a table you can take into a roadmap conversation.

The five verdicts

  • Pursue. It landed. Green.
  • Iterate. The core works, something around it doesn’t. Yellow.
  • Drop. It didn’t land. Red, and the cheapest result in the report.
  • Inconclusive. Not enough consistent signal to call it. Grey.
  • Rewrite-and-retest. Blue, because it isn’t a judgement on the concept at all. The description got in the way, so the test is unfinished.

Objection severity reads backwards

This is the single most misread thing in any Candor report, so it’s worth stating plainly. Severity describes how hard an objection is to answer, not how strongly participants voiced it.

  • Low severity, green. The concept has a strong answer to this objection. It’s addressable.
  • High severity, red. The concept has no good answer. The objection is structural.

So a wall of Low-severity objections is a healthy concept with a messaging job to do. A couple of High-severity ones is a concept that needs rethinking, however enthusiastic the rest of the report sounds.

What’s in Findings

The concepts one by one, sorted best to worst by verdict rather than by the order you entered them. Each row carries the verdict, a summary, the description, and a collapsed detail with the evidence. On a card-sort test (five or more concepts) the row also says how many participants were probed in depth on that concept.

Alongside them sit direct answers to the learning goals you set, and a section naming needs that none of your tested concepts covered. That last one is the closest a concept test gets to discovery, and it’s where the next idea usually comes from.

The action for each concept appears once, in Next Steps, rather than being repeated under every concept. A concept row that a step covers carries a See Step N link; clicking it opens Next Steps at the step. A concept no step names carries none.

What only appears with multiple concepts

Three things render only when the run used comparative mode, which means you tested two to four concepts: the head-to-head read on how participants chose, What would have to be true to adopt, and, for B2B audiences, Approval requirements covering who else has to sign off.

The middle one is the most practical section in the report, because it converts objections into conditions you could go and satisfy. Note the ceiling as well as the floor: a single concept is below comparative mode, and five or more switches the guide to a card sort, which does not produce these sections either. Two to four is the band that gets you the comparison.

In the head-to-head read, each line under What drove the choice carries a count taken from the interviews. A reason people gave whatever they chose reads like 4 of 12 participants. A reason tied to one option is counted only among the people who chose it, and reads like 5 of the participants who picked followed by the option. A reason with no count shows none.

Follow-on tests are labelled

If you ran a name, value proposition, or messaging test rather than a standard one, the report is titled by that sub-type, for example Value proposition test v2. Version numbers count within a sub-type, so your first name test is v1 even if standard concept tests already ran in that study. The report structure is the same; what changes is that the concepts being compared are options for one anchor concept rather than peer ideas.

Comparative reports count each participant's favourite out of everyone in the report, in the decision table's Favourite column, for example 18 of 24. On a name, value proposition or messaging test, and on a test run from an uploaded interview guide, the cell also says how many picks were reluctant, as in 18 of 24 (2 reluctant), as long as every pick on the report recorded it. A standard comparative test never shows a reluctant count. A reluctant pick means the participant said none of the options appealed, then named that one as their choice if they had to pick. Someone who names no option is not counted for any option. Reports made before this count existed show the older count, out of the picks alone.

Where to go next

Common questions

Good, and this is the one label in any Candor report most likely to be read backwards. Severity describes how hard the objection is to answer, not how loudly your synthetic participants raised it. Low severity in green means the concept already has a strong answer to that objection, so it's addressable with better wording or a feature you could add. High severity in red means the concept has no good answer, which makes the objection structural rather than a messaging problem. A concept with several High-severity objections needs rethinking, not rewriting.

Drop means the idea didn't land. Rewrite-and-retest means you can't tell yet, because the description got in the way. It's the only verdict coloured blue rather than green, yellow, or red, precisely because it isn't a judgement on the concept: it's a judgement on how the concept was written. Treat it as a signal that the test is unfinished. Rewrite the concept statement and run it again rather than reading it as a soft no.

Because the run wasn't in comparative mode, and there are two ways to miss it. One concept is below the threshold. Five or more switches the guide to a card sort, which ranks concepts but doesn't produce the head-to-head read, the What would have to be true to adopt section, or the B2B approval requirements. Comparative mode is two to four concepts, so that is the band to aim for when the point of the study is choosing between directions. Outside it you still get verdicts, comprehension, objections, and next steps per concept.

More FAQs →

Bring evidence to your next decision.

Start with a free project, or walk through Candor with us first.