CMO playbook & advanced tactics for A/B testing
A shipped false positive never announces itself. It goes into the codebase, into the quarterly deck as a 3% lift, and into next year's roadmap as a proven pattern. Nobody ever takes it back out. Microsoft's experimentation team has reported that only about a third of the ideas they test move the intended metric in the intended direction; the rest are flat or negative. Run a population like that through a 5% bar at 80% power and roughly one declared winner in ten is noise wearing a suit. Let teams stop tests when the line looks good and that share climbs fast. The visible symptom is a P&L that never reconciles with the sum of the lifts your team has claimed, and no single test that can be blamed for it.
The maths is not what fails at that point. What fails is who owns the platform, how fast a test is allowed to ship, and who has standing to veto a positive readout.
Core concept: experimentation as an institution, not an activity
Take the significance, power and sample size vocabulary as given from the foundations lesson. What that vocabulary does not tell you is who pays for the plumbing. An institution has four things a testing habit does not: one assignment service every surface calls, one library of metric definitions, an automated readout that runs identical checks on every test, and a decision log recording what shipped and why.
Microsoft's published work on experimentation maturity (Kohavi, Fabijan, Dmitriev and colleagues) describes an organisation moving from hand-built one-off tests to platform-served experimentation across Bing, Office, Windows and Xbox. The part that matters to a CMO is what changes between those stages. At the low end, an analyst is the platform, and analysis quality depends on which analyst picked up the ticket. At the high end an engineer ships a test without an analyst in the loop, because the checks are code and the metrics are already defined.
That capability is a standing engineering and data science cost, not a line item in a campaign budget. It also has a political consequence: if experimentation lives entirely inside marketing's budget, it gets cut in the first bad quarter, and the definitions of your own conversion metrics leave with it. Co-ownership with product and engineering is the arbitration worth fighting for early.
Sub-concept 1: build, buy, or stay small on purpose
Amazon built Weblab. Microsoft built ExP. Both did so because experimentation reaches into ranking, pricing and backend services, where a client-side vendor script cannot go. Vendors (Optimizely, VWO, Statsig, Eppo, GrowthBook, all of whom sell the thing under discussion) will run a competent platform for a fraction of that cost, and for most marketing organisations that is the correct answer. What you are buying is somebody else's statistics engine and assignment logic, which means you inherit their defaults.
The third option gets ignored. Take a brand whose best page sees 50,000 sessions a month at a 3% conversion rateconversion rateThe percentage of visitors or prospects who complete a desired action (purchase, sign-up, contact form), calculated as conversions divided by total opportunities.View full definition →. Detecting a 5% relative improvement needs roughly 200,000 sessions per arm, so about eight months of traffic for one test. No platform purchase changes that arithmetic. Below a certain traffic level the honest decision is to run few, large, structural tests, accept qualitative research and judgement for the rest, and stop pretending that a fortnight of data on a button is evidence. Buying a platform to run underpowered tests faster only industrialises the noise.
Sub-concept 2: velocity versus review quality
Two failure modes sit at opposite ends of the same dial. A weekly experiment review board staffed by senior people turns three tests a week into three a month, and the queue becomes the constraint on learning. No review at all, and a meaningful share of your readouts are broken before anyone reads them: mismatched traffic splits, a metric changed mid-flight, bots in one arm.
Amazon's 2013 shareholder letter reported 1,976 Weblab experiments run that year, against 1,092 the year before and 546 in 2011. Velocity of that shape is not achieved by reviewing harder. It comes from making the checks automatic and cheap: a sample ratio mismatch alarm on every test, a pre-registered primary metric and duration recorded before traffic starts, standing guardrailguardrailRules and controls that keep an AI system inside safe, legal and on-brand boundaries, blocking outputs and actions that cross the line.View full definition → metrics that fail a test regardless of what the primary metric says.
The peeking question belongs here too, as governance rather than statistics. Fixed-horizon tests read early inflate false positives badly; Optimizely, which sells an experimentation platform, published simulations showing continuous monitoring at a nominal 5% pushing real error rates far past 20%. A leader has two coherent positions: fixed horizon with early reading treated as a protocol violation, or an always-valid method so that reading early is legitimate. Choosing neither, which is the default in most organisations, means the number in the deck has no defined error rate at all.
Sub-concept 3: concurrency and interaction effects
Once dozens of tests run at once, they overlap on the same users. Microsoft runs large numbers of concurrent experiments and checks automatically for interactions between them rather than serialising the queue, which is the only affordable answer at scale. The risk is genuine: on a retail page, a headline change and an image change that each win alone can lose together, because a visitor reads the page as one object, not as independent slots.
Full factorial multivariate testing is where this gets expensive. Three variables at two levels each is eight cells, and eight cells need far more traffic than two. At Amazon's volumes that is a rounding error. For most brands it means the test finishes after the season it was meant to inform. Fractional factorial designs or sequential single-variable tests on the pages that carry real money are the practical route. The leadership decision is not whether to allow concurrency, it is whether your platform can detect an interaction when it happens, and whether tests touching the same surface are isolated by default.
How Booking.com Runs 1000+ Experiments
SUB-CONCEPT 4: WHAT COUNTS AS A WIN, AND WHO GETS TO SAY SO
Every mature programme converges on a single decision criterion agreed in advance, plus guardrails that can veto. Conversion rate alone is a bad criterion because it is easy to lift by borrowing from somewhere else: aggressive interstitials that raise signups and raise churn, a checkout that raises orders and raises returns. An Amazon experiment reported by Greg Linden found that every 100ms of added latency cost roughly 1% of sales, which is why page weight belongs on the guardrail list for anything a marketing team ships.
The strongest check on a portfolio of claimed wins is a long-run holdout: keep a small slice of users, 1% to 5%, out of everything you ship for a quarter or more, then compare their behaviour with the cumulative lift your test log claims. When the two disagree by a wide margin, you have a false discovery rate problem, and you have found it before your CFO does.
Real-world cases
Case 1: Microsoft Bing ad headlines
In 2012 a Bing engineer proposed a small change to how ad headlines were displayed. The idea sat in the backlog for months because it looked trivial. When it finally ran, it produced around a 12% increase in Bing's US revenue, over $100 million a year. Stefan Thomke's account in Harvard Business Review draws the organisational lesson: pre-filtering ideas by perceived importance is itself a costly decision, made by people with no way of knowing. The corollary is uncomfortable for a CMO. If your intake process ranks ideas by seniority of sponsor, you are paying for the ranking with the ideas you never test.
Case 2: Amazon's velocity as a target
Bezos has framed Amazon's success as a function of how many experiments it runs per year, per month, per week, per day, and the Weblab counts in the shareholder letters show that framing being managed as a number that roughly doubled year on year. Velocity as an explicit executive metric changes behaviour further down: teams stop bundling five changes into one release to save review time, because the review is not the bottleneck any more.
Case 3: Duolingo's compounding
Jorge Mazal's 2022 account of Duolingo's growth work describes a long run of small changes to streak mechanics, streak repair and notification timing, most of them worth a few percent at best. Duolingo reported daily actives above 16 million by the end of 2022, growing more than 50% year on year. No single test in that programme would survive an executive review asking for a business case. The programme did. That asymmetry is the argument for funding a testing capability rather than approving tests one by one.
A/B Testing at Scale: Lessons from Netflix
CMO action items
- Establish ownership in writing this quarter: name the team that owns assignment, metric definitions and readout tooling, and get the budget line co-signed with product or engineering rather than sitting alone in marketing.
- Pick your peeking regime and publish it. Either fixed horizon with locked duration, or an always-valid method. Then ask your team which one the current tool actually implements; the answer is often not the one they assume.
- Instrument the programme, not just the tests: experiments started per month, share failing the traffic-split check, share of readouts flat or negative, and time from readout to decision. A programme where 90% of tests win is measuring badly, not performing well.
- Run a long-run holdout on your highest-traffic surface for one quarter and reconcile it against the lift your test log claims for the same period.
- Require two segment cuts on every readout presented to leadership, new versus returning and mobile versus desktop, so that a flat average does not quietly kill a change that works for a large minority.
Common mistakes that kill results
Mistake 1: Making early stopping a career-positive act. Teams peek because someone senior asks how the test is doing on day four, and answering with a big number is rewarded. The fix is structural rather than educational: locked duration recorded before launch, and a readout template that shows the pre-registered end date next to the current result. If early stopping is treated as initiative, no amount of statistics training will stop it.
Mistake 2: Filling the roadmap with surface tweaks. Colours, spacing and font sizes are easy to test, easy to explain and rarely worth much. Pricing architecture, onboarding sequence, offer structure and value proposition copy are where the money is, and they are harder to design and riskier to run, which is why they slip. A roadmap that reads as a list of visual adjustments is a signal about your review process, not about your designers.
Mistake 3: Treating a shipped win as permanent. Effects decay as the market, the traffic mix and the competitive set move, and Microsoft's experimentation literature treats long-run effect measurement as a separate exercise from the two-week readout for exactly that reason. Put a re-test schedule on the handful of decisions that carry the most revenue, and accept that some of them will not replicate. Finding that out on purpose costs far less than discovering it through an unexplained quarter.
Resources
- 🔗Trustworthy Online Controlled Experiments by Ron Kohavi et al.
The definitive technical and strategic reference on A/B testing at scale, written by Microsoft's experimentation leaders with real case studies from Bing and other properties.
- 🔗Duolingo Growth Model: How A/B Testing Drove 20% DAU Growth
Jorge Mazal's published account of Duolingo's sequential testing program, including specific metrics on streak mechanics and notification experiments that compounded to 20% daily active user growth.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Lock test duration and sample size before launch; ban early peeking
- Prioritize testing high-stakes structural elements over low-traffic cosmetic tweaks
Related articles
Recent articles from the blog that build on this lesson.
- MarketingMeasuring influencer ROI past vanity metrics: a CMO's playbookLikes and follower counts tell you almost nothing about whether an influencer campaign moved your business forward. This playbook shows CMOs how to build a measurement framework that connects influencer spend to revenue, retention, and brand equity.
- MarketingBuilding an experimentation culture inside your marketing orgMost marketing teams run experiments occasionally. The ones that compound their advantages run them continuously, with infrastructure and incentives that make testing the default, not the exception.