A/B Testing
Also: Split Testing, A/B Test, Online Controlled Experiment
A/B testing is a controlled experiment that compares two versions of something (A and B) by splitting traffic randomly to learn which performs better on a chosen metric.
What It Is
A/B testing (also called split testing) is a controlled online experiment in which two or more variants are shown to randomly assigned groups of users at the same time. One group sees the control (version A, usually the current experience), and the other sees the variant (version B, the proposed change). By comparing a predefined metric between groups, you can attribute differences in behavior to the change itself rather than to chance or external factors.
The random assignment is the key. Because users are split randomly, the two groups are statistically similar in every other respect, so any meaningful difference in outcomes can be linked to the variant.
Why it matters
- Causal evidence: Unlike before/after comparisons, A/B testing isolates cause and effect, reducing the risk of acting on correlation or seasonal noise.
- Lower risk: Changes are validated on a fraction of traffic before a full rollout.
- Shared language: It gives marketing, product, and data teams an objective basis for decisions instead of opinion or hierarchy.
- Compounding gains: A steady cadence of small validated wins often outperforms occasional large redesigns.
How it is used in practice
1. Form a hypothesis: State the expected effect, for example "A shorter checkout form will increase completion rate."
2. Choose a primary metric: Pick one decision metric (conversion rate, revenue per visitor, click-through) plus guardrail metrics.
3. Compute sample size: Use the baseline rate, the minimum detectable effect, and a significance level to decide how long to run.
4. Randomize and run: Split traffic, run until the planned sample is reached, and avoid stopping early just because results look good.
5. Analyze: Check statistical significance and the confidence interval, then decide to ship, iterate, or discard.
Concrete Example
An e-commerce CMO tests a green "Buy now" button (B) against the current blue button (A). Traffic is split 50/50 across 40,000 visitors. Version A converts at 4.0 percent, version B at 4.6 percent. With a p-value below 0.05, the team concludes the lift is unlikely to be random and rolls out the green button to all users.
Common Pitfalls
- Peeking at results and stopping early inflates false positives.
- Too many metrics raise the chance of a spurious "winner."
- Underpowered tests that run too short cannot detect real effects.
See also
Frequently asked questions
What is A/B testing?
A/B testing, also called split testing, is a controlled online experiment where two or more versions of a page, email or feature are shown at the same time to randomly assigned groups of users. One group sees the control (version A, the current experience), the other sees the variant (version B). Because assignment is random, the two groups are comparable in every other respect, so a difference in the chosen metric can be attributed to the change itself.
What's the difference between an A/B test and a before/after comparison?
A before/after comparison measures the same audience at two different moments, so seasonality, a campaign launch or a news cycle can explain the change instead of your modification. An A/B test runs both versions at the same time on randomly split traffic, which removes those external factors. That is why A/B testing gives causal evidence while a before/after reading only gives a correlation.
Who in a company actually uses A/B testing?
Marketing, product and data teams, and the executives who arbitrate between them. For a CMO it validates a landing page, a pricing display or an email subject line before a full rollout; for a data leader it is the standard method for producing causal evidence. Its main organisational value is giving these teams an objective basis for decisions instead of opinion or seniority.
How do you decide how long to run an A/B test?
You compute the required sample size before launching, from three inputs: the baseline conversion rate, the minimum detectable effect you care about, and the significance level. The test then runs until that sample is reached, whatever the intermediate results look like. Deciding the duration afterwards, based on how the curves are moving, is what turns an experiment into a guess.
Why is stopping a test early when results look good a problem?
Because looking at results repeatedly and stopping at the first favourable moment inflates false positives: this is known as peeking. Metrics fluctuate naturally during a test, so any variant will cross the significance threshold at some point purely by chance. Two related pitfalls have the same effect: tracking too many metrics, which multiplies the chance of a spurious winner, and underpowered tests that run too short to detect a real effect.