Analytics

What is A/B Test

A controlled experiment comparing two versions to see which performs better.

Overview

An A/B test (split test) randomly divides your audience between a control version (A) and one variant (B), then measures which produces a better result on a chosen metric. Randomisation is what makes it powerful: it isolates the effect of the change from everything else going on, giving you causal evidence rather than correlation.

The test is only trustworthy if it is set up properly: a single clear hypothesis, a predefined primary metric, a sample large enough to detect a meaningful difference, and a run long enough to reach statistical significance before you decide.

How it works

Traffic is split randomly and simultaneously so both versions face the same conditions. You then compare the primary metric and check whether the difference is statistically significant or just noise.

  • One primary metric decided before the test
  • Adequate sample size and run length
  • Statistical significance before calling a winner
  • Randomised, concurrent exposure to remove bias

Common mistakes

Peeking at results and stopping the moment a variant looks ahead, testing during unrepresentative periods (a sale, a holiday), changing several elements so you cannot tell what caused the effect, or ignoring the false-positive risk of running many tests at once.

Why it matters

A/B testing is how you replace opinion with evidence: by showing two versions to comparable, randomly-split audiences simultaneously and measuring which performs better, you learn what actually moves your metric rather than what someone thinks will. That randomisation and simultaneity is what makes it causal — it controls for the seasonality, traffic-mix and external factors that make before/after comparisons unreliable. For any high-traffic page or funnel, it is the closest thing marketing has to a controlled experiment, and the antidote to the highest-paid-person's-opinion school of design.

  • Randomised, simultaneous split test — establishes cause, not correlation
  • Controls for seasonality and external factors that break before/after tests
  • Replaces opinion with evidence on what actually moves the metric

Getting valid results

The credibility of an A/B test rests on statistical discipline, and this is where most tests go wrong. You must define the sample size and duration in advance (based on your baseline rate and the effect you want to detect) and not peek-and-stop the moment a variant looks ahead — "peeking" and calling early is the number-one cause of false positives, because random noise routinely produces temporary "winners". Run tests for full business cycles (whole weeks) to avoid day-of-week bias, test one meaningful change at a time so you know what caused the result, and treat a non-significant test as a real, useful finding, not a failure.

  • Fix sample size and duration up front; do not stop the moment it looks significant
  • Peeking-and-stopping early is the top cause of false-positive "wins"
  • Run full weekly cycles to avoid day-of-week bias
  • A non-significant result is a valid, useful finding — not a failed test

In practice

A marketer launches an A/B test on a landing page and, two days in, sees the variant 'winning' by 20% with what the tool labels 95% confidence. Excited, they want to call it and roll it out. The disciplined response is to wait: the test was powered for a two-week run to reach an adequate sample across full business cycles, and stopping early ('peeking') is the number-one source of false positives, because random noise routinely produces temporary leaders that regress to the mean. They let it run, the gap narrows, and by the end the result is a genuine but smaller 4% lift — real, and worth shipping, but a quarter of the illusory early number. Had they called it early, they would have 'learned' a false lesson and been puzzled when it did not hold in production. The example is the cardinal rule of testing: pre-commit to sample size and duration, and never stop the moment it looks significant.

Common questions

A/B Test — questions

Straight answers on how this fits your marketing and build.

How long should an A/B test run?
Long enough to reach a predetermined sample size and statistical significance, and ideally covering full weekly cycles to avoid day-of-week bias. Stopping early because a variant looks ahead is the most common way tests mislead.
Why did my winning test not hold up after launch?
Usually the test was called too early, ran during an unusual period, or lacked the sample to be significant. That produces false positives that regress to the mean once the change is live for everyone.
Why did my "winning" A/B test not hold up after launch?
Almost always because it was called too early — you stopped the moment a variant looked ahead, capturing random noise rather than a real effect ("peeking"). Underpowered tests on too little traffic produce false winners routinely. Pre-commit to a sample size and duration and only judge the result at the end.
Can I test several changes at once?
You can, but a standard A/B test tells you the combined effect, not which change caused it. If you change the headline, the button and the image together and it wins, you do not know which mattered. To isolate individual effects you need multivariate testing (and much more traffic), or to test changes one at a time.

Still have questions? Talk to a specialist