Most advertisers test reactively: when a campaign stops working, they change the creative. Systematic creative testing reverses that process: you test continuously and proactively to always know which is the best creative available, before your current ones wear out.

The difference in results is consistent: advertisers with a systematic creative testing process maintain CPAs between 25% and 40% below those who react when results drop.

Why classic A/B testing is no longer enough

Classic A/B testing has a speed problem. To achieve statistical significance with 95% confidence in a two-variant test, you need between 500and 2.000 conversions per variant depending on the expected effect. With a volume of 100 conversions per month, each test takes between 10and 40months. By the time you have the result, the market context has changed.

The Bayesian model solves this problem. Rather than waiting for statistical significance, the model continuously updates the probabilities of each variant being the best, based on incoming data. From the first few hours of data, it can begin allocating more budget to the most promising variants, without waiting for a definitive winner to be declared.

This has a compound effect: winning variants receive more budget and more data sooner, which accelerates learning. And the budget wasted on clearly losing variants is minimised.

The process of generating variants at scale

Generating fifty variants of an ad does not mean creating fifty completely different ideas. It means systematically varying the five components of the ad:

Headline angle (10 variants): main benefit, urgency, curiosity, question, social proof, before/after contrast, negative (what to avoid), number, sensory experience, customer identity.

Body copy (5 variants): direct to benefit, story + resolution, problem + agitation + solution (PAS), social proof + CTA, objection + rebuttal + CTA.

CTA (5variants): "Start for free", "See how it works", "Book a call", "Download the guide", "Calculate your savings".

Visual (5variants): product photo, in-use photo, person + product, text only + brand colour, before/after.

Format (5variants by platform): static image 1:1,carousel, video 15s, video 30s, static image 4:5.

With AI, generating ten headline variants takes ten minutes, not two hours. The creative team defines the brief and the angles; the AI produces the variants; the team selects those that pass the quality filter. The result: 50 variants ready in a day instead of a week.

The phased testing structure

Phase 1 (weeks 1-2): Angle test. With a small budget, test which of the ten headline angles works best. The Bayesian model identifies the two or three winning angles within the first seven days.

Phase 2 (weeks 3-4): Depth test. For the winning angles, test all copy, CTA and visual variations. The budget is concentrated on the most promising combinations.

Phase 3 (weeks 5-8): Scale the winner. The winning combination receives the full budget. Testing of the next generation of variants begins straight away to have the replacement ready before the current winner burns out.

The system never stops: there is always one generation in the testing phase whilst the previous one is scaling. This ensures that when a creative fatigues (CTR starts to drop because the audience has seen it too many times), there is already a validated replacement ready to activate.

Our Creative Testing product implements this complete process with the Bayesian model and AI-generated variants.

Related reading