AdTest.AI
Blog

Concept testing: how it actually works

Concept testing is the step most advertisers skip and then pay for twice. It happens before the shoot, before the edit, before a single pound of media — while the idea is still a script, a storyboard or three lines in a deck. The point is not to find out whether an ad works. It is to find out which idea deserves to become an ad. Get that decision right and everything downstream gets cheaper; get it wrong and no amount of optimisation rescues it.

What concept testing actually is

A concept is the idea stripped of production: the promise, the situation, the reason to care. Concept testing evaluates that idea against alternatives, with the finish deliberately held constant so that craft does not contaminate the comparison. It is distinct from finished-ad testing, which asks whether this execution works, and from copy testing, which interrogates the words once the idea is settled.

The practical consequence is a sequencing rule: test concepts to choose a direction, test executions to choose a cut, and test in-market to allocate spend. Teams that collapse these into one stage end up A/B testing two executions of an idea that was never the strongest of the five they had.

The three designs, and when each is right

  • Monadic — each respondent sees one concept and rates it in isolation. The cleanest read on absolute appeal, and the closest analogue to how anyone meets an ad in the wild. It is also the most expensive, because every concept needs its own sample.
  • Sequential monadic — each respondent sees several concepts in rotated order, rating each before moving on. Far more efficient, and the default for most commercial work. Order effects and fatigue are real, which is why rotation is not optional and four to five concepts is a sensible ceiling.
  • Comparative — respondents see concepts side by side and choose. Excellent at ranking, poor at telling you whether the winner is any good. A comparative test will always produce a champion, including from a set where every option is weak.

The failure mode is picking comparative because it is cheap, then reporting the winner as though it had cleared a bar. It cleared a field.

What to measure at concept stage

Concept-stage diagnostics should predict potential, not polish. Five carry most of the weight:

  • Distinctiveness — is there anything here a competitor could not have said? This is the single best early predictor and the one most likely to be sanded off in later rounds.
  • Comprehension — can someone play the idea back after one exposure? If the playback is vague, the media plan will not fix it.
  • Relevance — does it land against a real tension for the audience, or a marketing-plan tension?
  • Motivation — does it move intent, and does the claimed benefit survive contact with scepticism?
  • Brand linkage — would the audience attribute it to you, or to the category leader? Weak linkage is how a well-liked concept ends up advertising someone else.

These map onto the broader framework we use in the 13 dimensions of creative effectiveness, with the production-dependent dimensions held back until there is something produced to judge.

Where traditional concept testing breaks

The method is sound; the delivery is where it fails. Panel-based concept tests typically take one to three weeks and cost thousands per wave, which means most teams run one wave and treat it as final. That has three consequences. Iteration stops, because a second wave costs as much as the first. Sample sizes get squeezed, so genuinely small differences between concepts are read as signal when they are noise. And respondents rationalise: asked why they preferred a concept, people produce articulate reasons that have little to do with the choice they made.

There is also a selection problem nobody likes discussing. Concepts that reach testing have already survived internal politics. The test adjudicates between three survivors, not between the ten ideas that existed on Monday.

A workflow that fits how creative actually gets made

The version that works in practice is two-stage. Score every concept on the table — all ten, not the three that survived the meeting — with an AI model that returns a structured read in minutes rather than weeks. Kill the bottom half on evidence rather than seniority, sharpen the survivors against their weakest diagnostic, and rescore. Then put human research behind the two or three that are left, where the sample size and the cost are actually justified.

That is the argument for mixing machine scores with human reads: the machine is fast enough to make iteration free, and the humans are reserved for the decision that carries the budget. It is also the fastest way to stop wasting your testing budget on measuring things you have already committed to. AdTest.AI scores a concept in about ninety seconds.

Frequently asked questions

What is concept testing in advertising?

It is the evaluation of an advertising idea before production — usually as a script, storyboard, or written proposition — to decide which idea should be made. It differs from ad testing, which evaluates a finished execution.

How many concepts should you test at once?

Four to five in a sequential monadic design. Beyond that, respondent fatigue degrades the later ratings even with rotation, and the diagnostics get noisier exactly where you need them sharpest.

What sample size does a concept test need?

Traditional panel work targets roughly 100 to 150 respondents per concept for stable diagnostics. Smaller samples still rank concepts, but they cannot reliably separate two that finish close together — which is usually the decision you are trying to make.

Is concept testing the same as copy testing?

No. Concept testing chooses the idea before production; copy testing evaluates the language and messaging of an execution. Concept testing comes first, and a good concept test narrows what copy testing has to cover.

Can concept testing be done without a panel?

Yes. AI-based scoring reads concepts against effectiveness criteria in minutes and at negligible marginal cost, which makes it practical to test everything rather than a shortlist. Panels remain worth their price for validating the final one or two.

Share this:

More from Blog

Incrementality testing: what it can prove

Incrementality testing is the only method that answers the question every other measurement tool dodges: would this conversion have happened anyway? Attribution tells you which touchpoint was nearest…

Read more

Brand lift studies: what they cost

A brand lift study is the measurement brand teams trust most and understand least. It carries the authority of a controlled experiment, it produces a number a CMO…

Read more

Ad fatigue: how to catch it early

Ad fatigue is the point where a creative stops earning its impressions. Nothing about the targeting changed, the budget is the same, the offer is the same —…

Read more

Accessibility Toolbar