+75 XP

Real-world application of A/B testing

Booking.com's internal experiment tool has a screen most marketing teams never build: everything running right now, on the order of a thousand experiments at once, each with a named owner, a pre-committed stop date, and guardrail metrics that can end the test without anyone asking permission. The company's published position, set out in its 2017 paper "Democratizing online controlled experiments at Booking.com", is that essentially every change ships behind a flag as an experiment. This lesson stays inside that one programme and reads its output the way the people there do, which means spending most of the time on the readouts that were stopped, flat, or right about the metric and wrong about the business.

Anatomy of a readout

A readout at Booking.com is not a single number. Significance, power and the required sample are settled before launch, on the terms the foundations lesson sets out, and the split runs on the persistent visitor profile described there. What arrives at the end is a table: the primary metric with its interval, a row of guardrails (page latency, error rate, cancellations, customer service contacts), the segments that were declared in advance, and a decision field with three real options: ship, kill, iterate.

The decision field is the interesting part. Roughly one tested idea in ten produces the improvement it was designed to produce. Lukas Vermeer, who ran experimentation at Booking.com, has repeated that order of magnitude publicly for years, and it matches what Microsoft and others report from their own programmes. If nine readouts in ten end in kill or iterate, the capability you are buying is not the ability to recognise a win. It is the ability to close a test fast, cheaply, without a meeting, and still write down one thing you did not know before.

Four things a readout must survive

  1. The guardrail row, read before the primary metric

The primary metric answers the question you asked. The guardrails catch the question you did not. A variant that lifts bookings while pushing customer service contacts up has moved cost, not value, and the conversion line alone will never show it. At Booking.com scale a guardrail breach triggers a stop rather than a debate, which is why the stop date can be pre-committed at all.

  1. The economics of a 90% failure rate

If nine tests in ten fail, the programme only pays if a test costs roughly what writing the code once costs. The moment each experiment carries a week of engineering setup, a bespoke dashboard and a review panel, the arithmetic collapses and the organisation starts protecting ideas instead of killing them. Killing early also frees the scarcest input: traffic.

  1. What your volume can actually see

Booking Holdings reports on the order of a billion room nights a year, and even that does not make small effects free. A half-percent relative move on conversion still takes weeks in a single market, and a thousand concurrent experiments are all drawing from the same visitors. Prioritisation is a traffic budget, not a wish list. Deciding to test the search results page is deciding not to test something else this month.

  1. The lag between exposure and truth

Travel has a long tail: exposure today, booking next week, stay in three months, cancellation possible at any point in between. A variant that pressures people into bookings they later abandon will move the primary metric and destroy value. Nothing catches that except a cancellation guardrail and a measurement window long enough to include the outcome, plus at least two full weekly cycles to absorb day-of-week patterns and novelty.

How Booking.com Uses A/B Testing at Scale

Watch on YouTube

Three readouts: stopped, flat, and wrong

Case 1: the experiment designed to lose

Booking.com engineers deliberately injected latency into the site to price what speed is worth. In the results published in their 2019 KDD paper, roughly a 30% increase in latency cost about 0.5% of conversion. The test was run to be stopped; its value is the exchange rate it produces. Any later feature that wins 0.3% on conversion while making the page materially heavier is now a known net loss, and the argument about whether it "feels" faster is over. Most companies never run a negative experiment and therefore have no price for the thing they trade away every sprint.

Case 2: 150 models, and the metric that lied

The same paper covers 150 machine learning models put into live experiments. One finding is uncomfortable: gains in offline model performance did not reliably translate into business value, and in a number of cases models that scored better offline performed worse in the experiment. The readouts came back flat or negative on the metric that pays. The offline score was a proxy, the proxy drifted from the outcome, and only the controlled test exposed the gap. If your agency reports model accuracy, engagement scores or brand lift as evidence of value, this is the case to hold them against.

Case 3: right on conversion, wrong on everything else

Booking.com's urgency and scarcity messaging (rooms left, other people viewing, discount framing) is the most copied pattern in online travel, and it is copied because it wins on the conversion metric. In February 2019 the UK Competition and Markets Authority obtained formal undertakings from Booking.com and five other travel sites covering pressure selling, discount claims, hidden charges and how search ranking is disclosed. The site changed. No readout could have shown that, because no primary metric and no guardrail in the tool measures a regulator, a journalist or a customer who books once and never trusts the brand again. Experimentation optimises inside the boundary you draw; drawing the boundary is a human decision and it belongs before the test, in writing.

A/B Testing Statistics Explained Clearly

Watch on YouTube

CMO action items

  • Ask for the kill rate, not the win rate. A team reporting that most of its tests won is either testing trivialities or reading the data badly. If the number is not somewhere near nine in ten not shipping, find out why before you celebrate.
  • Require a guardrail row on every readout you are shown, with latency and a downstream quality metric (cancellations, refunds, returns, support contacts) on it. A readout with one line on it is a sales pitch.
  • Write the decision rule into the test document before launch: what result ships, what result kills, what result triggers a rerun, and the date. Then hold the date.
  • Keep the losses. A results library that records the hypothesis, the readout and the estimated value of the ones that failed stops your organisation from re-testing the same dead idea every eighteen months as staff turn over.

Common mistakes that kill results

  • Peeking and stopping on the first good day. Watching a live test and calling it the moment the line goes green inflates false positives badly. Either commit to the end date or use a sequential method built for continuous monitoring, and say which one you are using before launch.
  • Hunting for a segment after the fact. Slice a flat result by device, market, traffic source and new versus returning and you will find a "winner" by chance alone; twenty slices at the usual threshold produce one false positive on average. Declare the segments you care about in advance, and treat anything else as a hypothesis for the next test rather than a result.
  • Assuming your concurrent tests are independent. When a hundred experiments touch the same funnel, two variants can interact, and a shared traffic pool means each one is also diluting the others. Large programmes handle this with isolation groups and interaction checks. A team running twelve tests on one landing page with none of that is generating noise.
  • Importing someone else's winner. Booking.com's results hold for Booking.com's traffic, price points and intent. A pattern that lifts conversion for a customer choosing a hotel for next Friday tells you close to nothing about a six-month enterprise sales cycle. Borrow the hypothesis, never the conclusion.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Lock test duration and sample size before launch; ban early peeking
  • Prioritize testing high-stakes structural elements over low-traffic cosmetic tweaks
  • Maintain a shared test results library documenting every win and loss
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.