A/B testing & statistics: foundations & core concepts
Split a mailing list in half and send both halves the identical email. The two open rates will differ. They always differ. That gap is noise, and it is the reason experimentation needs a vocabulary instead of an argument.
Microsoft's experimentation team, which has run controlled tests across Bing and Office for well over a decade, has reported that only about a third of the ideas that reach a test improve the metric they were built to improve. Some do nothing. Some make things worse. Take that seriously and an uncomfortable conclusion follows: if most ideas fail, then most celebrated "wins" from tests stopped early or run on thin traffic are randomness with a green arrow next to it. The words in this lesson (hypothesis, randomisation unit, significance, power, sample size) exist to tell the two apart.
What a/b testinga/b testingA/B testing is a controlled experiment that compares two versions of something (A and B) by splitting traffic randomly to learn which performs better on a chosen metric.View full definition → actually is
An A/B test, also called a split test or an online controlled experiment, splits your audience at random into two groups. Group A sees the current version (the control). Group B sees the changed version (the variant). Both run at the same time, on the same traffic mix, and you compare one outcome declared in advance.
Random assignment is what makes the comparison legitimate. Because the split is random, the groups differ only by chance and by the change you made, so a gap bigger than chance can be attributed to the change. Compare this week to last week instead and you get a before/after story polluted by weather, payday, a competitor's promotion and your own media spend.
The randomisation unit is the thing you split. Usually the visitor (a cookie or a logged-in ID), sometimes the session, occasionally a city or a store. Split by visitor and someone who returns four times always sees the same version, which is what you want when the change affects how people judge a brand. Split by session and the same person can see both versions, which both corrupts the comparison and confuses the customer. The unit also has to match the metric: revenue per visitor needs visitor-level randomisation.
Four statistical ideas carry the rest.
The null hypothesis is your starting assumption: there is no difference between control and variant. A test never proves the variant is better. It collects enough evidence to make "no difference" implausible.
Statistical significance is where you set the bar for implausible. The convention is 95%, and its meaning is narrow: if the null hypothesis were true, only 5% of tests like yours would show a difference this large.
The p-value is what you compare against that bar. It is the probability of seeing a result at least this extreme if there were genuinely no difference. Below 0.05, the result is called significant. It is not the probability that your variant works, and that confusion is where most misreadings start.
Power is the idea marketers skip, and skipping it is why their results are noisy. Power is the chance your test detects a real effect of a given size. 80% is the usual target. A test with 30% power will miss most genuine improvements, and the few "wins" it does report come back inflated or in the wrong direction.
Key sub-concepts every CMO must own
Sample size: the most abused number in marketing experimentation. It is not a duration you choose, it is an output. Feed a calculator (Evan Miller's is free at evanmiller.org) your baseline conversion rateconversion rateThe percentage of visitors or prospects who complete a desired action (purchase, sign-up, contact form), calculated as conversions divided by total opportunities.View full definition →, the effect you want to detect, your significance level and your power, and it returns the visitors needed per variant. If the answer is 400,000 per side and you get 20,000 a month, you have learned something useful: this test cannot be run here, and you should test something with a larger expected effect.
Minimum detectable effect (MDE): the smallest improvement you would act on. Detecting a move from 3.2% to 3.4% demands an enormous sample. Deciding you only care about relative lifts of 10% or more shrinks the requirement sharply. Set the MDE first. Choosing it after seeing the data is how teams talk themselves into a result.
Novelty and primacy effects: regular users react to a change because it is new, not because it is better. Early numbers move, then fade. Split results between new and returning visitors. If the lift lives only with returning users in week one, you are measuring surprise.
SegmentsSegmentsDividing a market into distinct groups of customers who share similar needs, characteristics or behaviours, so each group can be served with a tailored approach.View full definition →: a change can win on mobile and lose on desktop. Zalando, selling across more than 20 European markets, faces this constantly, since a checkout change can behave differently by country, by device and depending on whether the visitor arrived during a sale. Cut your results by the two or three segments that matter before shipping everywhere, but treat every extra cut as another comparison rather than another discovery.
How Obama Raised $60 Million by Running a Simple Experiment
Real company examples with real numbers
Microsoft Bing, 2012. An engineer proposed a change to how ad headlines were displayed. It sat in the backlog for months because nobody believed it mattered. When it ran, revenue per search rose roughly 12% in the US, worth over $100 million a year, and the team's first reaction was to suspect a tracking bug. Two lessons: intuition cannot rank ideas, and an extreme result earns an instrumentation check before a celebration.
Netflix has published engineering write-ups on artwork testing, showing that the image chosen to represent a title changes how often members play it. The same company also puts serious work into the statistics behind those readouts, including sequential testing methods that allow results to be monitored continuously without the false positive inflation that casual peeking causes. That is the tell of a mature programme: the method gets as much attention as the idea.
Zalando built its own experimentation platform rather than buying one, and its engineers have described the plumbing publicly: assigning a visitor to a variant consistently across visits, and computing metrics on the same unit used for the split. Unglamorous work, and it is what stops a fashion retailer from reading a two-week test that straddles a sale period as a permanent effect.
Knowledge check
1. What is the fundamental purpose of running an A/B test in marketing?
2. In an A/B test, what does the null hypothesis represent?
3. A marketer runs a test for three days, sees the variant winning by 8%, and declares victory. Why is this reasoning flawed?
4. Select ALL correct statements about p-value and statistical significance.
Select all the correct answers.
5. Select ALL inputs you typically provide to a sample size calculator before starting an A/B test.
Select all the correct answers.
CMO action items
- Insist every test arrives as one written sentence: because [evidence], we expect [change] to move [primary metric] by at least [MDE], measured on [randomisation unit]. A request that cannot fill in those brackets is a preference, not a hypothesis.
- Pre-register the numbers before launch: baseline rate, MDE, significance level, power, the resulting sample size per variant, the end date, and exactly one primary metric. Secondary metrics are for diagnosis, never for declaring the winner.
- Ask your tool what it actually computes. Google Optimize shut down in September 2023, so most teams now sit on Optimizely, VWO (both vendors that sell testing software, so read their benchmarks with that in mind), GA4 experiments or something in-house. Find out whether yours assumes a fixed sample size or supports sequential analysis, because that single answer decides whether daily checking is safe.
- Convert your MDE into money before running anything. A 10% lift on a page 300 people see per month is not worth an engineering sprint.
Common mistakes that kill results
Peeking and stopping early. Checking daily and stopping the moment the dashboard turns green does not accelerate learning, it manufactures false positives: research by Ramesh Johari and colleagues at Stanford on continuous monitoring showed the true error rate rising well above the nominal 5% when tests are read repeatedly with fixed-horizon maths. You agreed a sample size. Honour it, or move to a method built for continuous reading.
Testing many things at once and calling it an A/B test. Change the headline, the hero image and the button together and you cannot say which one did the work, or whether one gain hid one loss. Test one variable, or use a multivariate design that isolates each contribution. Optimizely and VWO both support this. The same arithmetic applies to metrics and segments: check 20 comparisons at 95% and about one will look significant by chance alone.
Confusing statistical significance with business significance. A 0.3% lift in click-through with a p-value of 0.03 is real and can still be worthless if shipping it costs $50,000 of engineering time. Translate every result into projected revenue or cost before the shipping decision.
Key takeaways
- An A/B test compares a control and a variant on randomly assigned, simultaneous traffic, judged on one outcome named in advance.
- Pick the randomisation unit deliberately (usually the visitor) and match your metric to it.
- Significance tells you how surprising your result would be if nothing were happening; power tells you whether your test could have found the effect at all. Underpowered tests are the main source of fake wins.
- Sample size and MDE are inputs you calculate before launch, not durations you choose afterwards.
- Roughly a third of tested ideas at Microsoft improved their target metric, so expect most of your hypotheses to fail and design tests that can survive that truth.
Resources
- 🔗Evan Miller's Sample Size Calculator
A free, no-login-required tool that calculates exactly how many conversions per variant you need before your A/B test results are statistically trustworthy.
- 🔗Trustworthy Online Controlled Experiments by Kohavi, Tang & Xu
The definitive book on running A/B tests at scale, written by the people who built experimentation programs at Microsoft, Google, and Amazon.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Lock test duration and sample size before launch; ban early peeking
- Maintain a shared test results library documenting every win and loss
Related articles
Recent articles from the blog that build on this lesson.