A controlled experiment comparing two versions to see which performs better.
An A/B test (split test) randomly divides your audience between a control version (A) and one variant (B), then measures which produces a better result on a chosen metric. Randomisation is what makes it powerful: it isolates the effect of the change from everything else going on, giving you causal evidence rather than correlation.
The test is only trustworthy if it is set up properly: a single clear hypothesis, a predefined primary metric, a sample large enough to detect a meaningful difference, and a run long enough to reach statistical significance before you decide.
Traffic is split randomly and simultaneously so both versions face the same conditions. You then compare the primary metric and check whether the difference is statistically significant or just noise.
Peeking at results and stopping the moment a variant looks ahead, testing during unrepresentative periods (a sale, a holiday), changing several elements so you cannot tell what caused the effect, or ignoring the false-positive risk of running many tests at once.
A/B testing is how you replace opinion with evidence: by showing two versions to comparable, randomly-split audiences simultaneously and measuring which performs better, you learn what actually moves your metric rather than what someone thinks will. That randomisation and simultaneity is what makes it causal — it controls for the seasonality, traffic-mix and external factors that make before/after comparisons unreliable. For any high-traffic page or funnel, it is the closest thing marketing has to a controlled experiment, and the antidote to the highest-paid-person's-opinion school of design.
The credibility of an A/B test rests on statistical discipline, and this is where most tests go wrong. You must define the sample size and duration in advance (based on your baseline rate and the effect you want to detect) and not peek-and-stop the moment a variant looks ahead — "peeking" and calling early is the number-one cause of false positives, because random noise routinely produces temporary "winners". Run tests for full business cycles (whole weeks) to avoid day-of-week bias, test one meaningful change at a time so you know what caused the result, and treat a non-significant test as a real, useful finding, not a failure.
A marketer launches an A/B test on a landing page and, two days in, sees the variant 'winning' by 20% with what the tool labels 95% confidence. Excited, they want to call it and roll it out. The disciplined response is to wait: the test was powered for a two-week run to reach an adequate sample across full business cycles, and stopping early ('peeking') is the number-one source of false positives, because random noise routinely produces temporary leaders that regress to the mean. They let it run, the gap narrows, and by the end the result is a genuine but smaller 4% lift — real, and worth shipping, but a quarter of the illusory early number. Had they called it early, they would have 'learned' a false lesson and been puzzled when it did not hold in production. The example is the cardinal rule of testing: pre-commit to sample size and duration, and never stop the moment it looks significant.
Part of our defined terms knowledge graph — browse every entry in this branch.
A structured process for increasing the share of visitors who convert.
The share of visitors who complete a desired action.
A measure of how actively people interact with content or a page.
The visible, clickable words in a hyperlink, used as a relevance signal.
A defined contract that lets two software systems talk to each other in a predictable way.
Common questions
Straight answers on how this fits your marketing and build.
Still have questions? Talk to a specialist