If you run consumer insights at a CPG brand, the economics of concept testing have quietly broken. Volume keeps rising and panel budget keeps not. The portfolio that needs testing is bigger every year, and panel rounds still cost $15K to $50K each (industry median around $23K) and take four to eight weeks. Most concepts get killed in a brand-team meeting before any respondent sees them. Some of those concepts were probably winners. You just don't get to know.
This essay is the argument for putting synthetic research first in CPG concept testing. Not synthetic-only. Synthetic-first. Synthetic screens the whole portfolio. Panel research concentrates on the survivors that earned the rigor. This is the routing decision most insights teams will end up making this decade, and the teams that make it sooner will get more research done at a similar or lower total spend.
What "synthetic-first" actually means
The phrase has a specific meaning. Synthetic-first is not "skip the panel." It is a sequencing decision about where each concept enters the research stack.
In the traditional sequence, every concept that reaches research goes straight to panel. The brand team's selection happens before research, by hand, in a room. Whatever survives the room gets the $23K panel round. Whatever doesn't survive the room never gets tested at all.
In the synthetic-first sequence, every concept that reaches research starts in synthetic. The full project pool (typically 10 to 50 concepts) gets respondent-grounded reactions in roughly one to two hours of platform time per concept. The brand team's selection happens after the first respondent-style signal, not before it. The concepts that survive synthetic advance to panel research with documented reasoning behind their survival, not just brand-team conviction. The panel rounds become richer because they're studying validated concepts in depth instead of spreading thin across un-screened candidates.
The total panel spend doesn't necessarily change. What changes is what the panel spend buys. Instead of $138K to $300K of panel research across six rounds that screen and validate at the same time, you get $69K to $200K of panel research across three to four rounds that validate concepts already screened. Less waste. Deeper studies. Better launch-gate decisions.
Why CPG specifically
The math hits CPG harder than most categories for three reasons.
Volume. Big brands now test 10 to 50 concepts per project across line extensions, packaging variants, claim variations, campaign angles, channel-specific positioning, and DTC sub-brands. Most consumer-insights teams run multiple such projects a year, reaching hundreds of concepts tested in some form across the portfolio. The traditional concept-testing cycle was built for a slower product world where a brand launched a few large bets per year and each bet got rigorous panel research. The volume has moved. The panel budget hasn't.
Per-round cost. Panel concept testing in CPG sits at the higher end of the industry-median curve. Beauty, food and beverage, and household-care concept rounds often run $30K to $50K because the categories demand higher-quality recruitment and longer survey instruments. Beauty and apparel concepts that require image-rich stimuli push even higher. Across the portfolio, that scales fast.
Speed of iteration. Brand teams iterate concepts continuously. The first synthetic round informs concept revisions; revised concepts go back through synthetic for re-test; only the survivors of two or three synthetic iterations reach panel. The iteration cost in traditional concept testing is the round cost: every revision is a fresh $23K panel commitment. The iteration cost in synthetic research is the platform time, which is hours.
The combination (high concept volume, high per-round cost, high iteration frequency) is the worst-case for traditional concept testing and the best-case for synthetic-first sequencing.
What synthetic does well in CPG concept testing
These are the concrete jobs synthetic does well, in CPG specifically.
Portfolio-wide concept screening. Run every concept in the project pool through synthetic concept testing. Each concept gets reactions from 8 to 16 evidence-grounded personas with calibrated personality and bias profiles. The output is qualitative-style depth on which concepts resonate, which fall flat, and the reasoning behind both. The brand team gets respondent-style reasoning to kill weak concepts, not just internal politics.
Claim screening. A new claim under consideration ("clean ingredients," "clinically tested," "30-day money-back guarantee," whichever the category supports) gets reactions from a population of category buyers across the segments you care about. Which claim angle resonates most. Which gets distorted into a meaning you didn't intend. Which feels overstated. Synthetic research handles relative-comparison claim screening well. It does not handle claim substantiation. The synthetic layer identifies the strongest candidates. The panel layer substantiates the survivors with the rigor a regulatory submission requires.
Price-anchor exploration. Across a portfolio of concepts at different positioning tiers, synthetic price testing surfaces where price-sensitivity breaks for each persona, without the panelist contamination that comes from respondents who've seen too many price tests in the past quarter. You get directional signal on tier structure, anchor sensitivity, and bundle preference before committing panel budget to a price round.
Discovery for category-entry concepts. A new sub-brand, a category extension into adjacent shelves, a DTC-only launch line. Traditional concept testing struggles here because the relevant population is hard to recruit (cross-category buyers, lapsed-purchase reactivation, new-buyer segments). Synthetic personas grounded in published consumer research can surface initial directional signal in hours rather than the months of recruitment a niche-population panel round demands.
Iteration between panel rounds. A panel round delivers findings. Concepts get revised. Without synthetic, the revised concepts either go straight back to panel ($23K + 4 to 8 weeks) or get advanced on judgment. With synthetic, revised concepts get pressure-tested against the persona pool in hours, and the panel round you're saving for the next gate gets concepts that have already survived two or three iterations.
What synthetic does not do in CPG, and why panel research remains required
The synthetic-first argument depends on being honest about what synthetic doesn't do. Insights leaders will notice immediately if we hide it.
Statistical point estimates with confidence intervals. "27% of category buyers prefer Concept A, plus or minus 3 percentage points at 95% confidence" requires real-respondent N drawn with documented sampling methodology. Synthetic research produces directional signal across persona variance, not statistically-bounded point estimates. If your launch committee or retail buyer demands a specific percentage with a specific confidence interval, that's a panel job by definition.
Substantiated claims. Anything going on a pack, in an ad, in a clinical-claim document, or in a regulatory filing needs real-respondent data with documented methodology. Synthetic research is appropriate for exploring claim angles and screening which ones are worth substantiating. It is not appropriate for the substantiation itself. The legal exposure is real. Stay on panel methodology where the regulator expects panel methodology.
Established category benchmarks. CPG concept testing has decades of category-norm benchmarks (volumetric purchase intent thresholds, top-two-box norms by sector, action-standard pass rates) that exist because they were built on panel data. New methodologies inherit category trust slowly. If your gate decision depends on comparing the candidate concept to a category norm with decades of panel-based history, that comparison happens on panel, not synthetic.
Brand health and tracking. Wave-over-wave consistency, year-over-year benchmarks, real-perception drift in-market. These depend on real-respondent panels sampled consistently across time. Synthetic research is moment-in-time and isn't designed for tracking studies.
Anonymized aggregation at scale. Segmentation work with 5,000 responses, share-of-wallet analysis, category-share studies. Panel infrastructure handles that scale, and synthetic research is a different tool entirely.
The rule that holds across all five: synthetic research is appropriate for the screening and iteration phases of concept work. Panel research is required for the validation, substantiation, and benchmarking phases. The two are complements, not substitutes.
The agreement-bias acknowledgment
CPG insights leaders are sophisticated about respondent bias. If a piece on synthetic research in CPG doesn't address agreement bias directly, the leader stops trusting the source. So this section directly.
Synthetic personas built on large language models exhibit a documented agreement bias. They tend to be more cooperative with the framing of a question than real respondents would be. A concept presented in flattering terms can score better in synthetic than the same concept would in panel research because the persona is more willing to engage with the framing the team is using to describe it. This is a real limitation, and ignoring it would be dishonest.
The way to use synthetic research given this limitation:
- Anchor on relative comparison, not absolute magnitude. "Which of these five concepts is strongest" is a question synthetic handles well. "How well will this concept perform" is a question synthetic handles poorly. Compare concepts against each other, not against an absolute reference.
- Use calibrated bias modeling. Candor's bias library treats acquiescence bias as a modeled trait, not an unmodeled assumption. The bias isn't eliminated, but it's made explicit and partially controllable rather than hidden.
- Use critic validation. A separate agent checks each response against the persona's established profile for consistency, which catches some of the agreement-bias artifacts before they show up in the synthesis output.
- Cross-check with panel data when the question matters. The synthetic layer surfaces directional signal. The panel layer validates magnitude. When the concept is close to a launch gate, both layers run, and the panel layer is the binding source of truth on magnitude.
The honest framing: synthetic concept testing is best at relative comparison, claim screening, and discovery. It is worst at absolute magnitude and at single-concept "will this work" questions. Use it for what it's good at and don't ask it to be what it isn't.
What synthetic-first looks like in a typical launch cycle
The pattern in CPG brands running synthetic-first across a typical new-product launch cycle:
Months 1 to 3. Brand team generates the concept pool (this is the slow step, not the testing). All concepts go through synthetic concept testing in Candor: a single focused work week of insights-team time across the portfolio. The brand team reviews the synthetic output and kills the unpromising concepts, advancing six or so survivors to deeper work. Each surviving concept gets value-prop and price-sensitivity exploration in synthetic in another work week. Concepts get sharpened, repositioned, or retired based on respondent-style reasoning, not internal politics.
Months 4 to 5. Three to four finalists go to panel concept testing. The rounds are richer because they're studying validated concepts in depth instead of spreading across the un-screened portfolio. Statistical confidence, category-norm benchmarking, and claim substantiation happen here.
Months 6 to 9. Production, shelf-ready package design, retailer presentations, brand health pre-launch measurement, advertising development. Synthetic research can support value-prop iteration and message testing during this window for the marketing team. Panel research supports advertising substantiation and pre-launch tracking.
Month 10 onward. Launch. In-market behavior, sales-pull data, real customer tracking. Panel research takes over from synthetic completely.
A consumer-insights team running this pattern moves multiple product launches per year through structured discovery-to-launch research, with synthetic doing the screening and iteration work in months 1 to 3 and panel research concentrated where it matters most.
The honest closing
Synthetic-first is the right routing decision for CPG concept testing today. It is not the right routing decision because synthetic research is better than panel research. It is the right decision because synthetic research and panel research do different things, and the sequence "synthetic for screening and iteration, panel for validation and substantiation" gets more research done across the portfolio than "panel for everything that survives an unobserved brand-team meeting."
The insights teams that adopt synthetic-first earliest will get more concepts tested, kill more concepts on respondent-style reasoning, validate fewer but stronger concepts on panel, and concentrate methodology rigor where the launch decision actually requires it. That's a stronger insights operation than the alternative, which is the operation most teams are running today by default rather than by choice.
The case for synthetic-first isn't a case against panel research. It's a case for using panel research on the questions panel research was built for, instead of using it as the only respondent-grounded research a brand team can afford to run.
If you're considering synthetic-first sequencing, the prior essay on when concept testing with synthetic users works covers the framework for routing individual concepts to the right layer. The methodology essay on why synthetic research needs evidence grounding covers the underlying methodology that makes synthetic research a research instrument rather than AI improvisation. Both are prerequisites to running this sequencing well.
For the broader head-to-head between synthetic and panel research, see Candor vs traditional research panels. For the role-specific walkthrough, see Candor for consumer insights teams.