+75 XP

Real-world application of GRP, TRP & offline metrics

On 4 February 2010 Old Spice put a 30-second spot online, and over Super Bowl weekend Wieden+Kennedy's "The Man Your Man Could Smell Like" ran on network television without an in-game buy. Within a week P&G had two accounts of what had happened and they did not agree. One was panel-reported audience delivery from Nielsen. The other was a spike in brand search, YouTube views and traffic to Old Spice properties that no panel had measured. Getting from two accounts to one verdict is the actual work here, so this lesson stays with that single campaign instead of touring five. Everything below assumes the currency itself: impacts, GRP against TRP, reach and frequency as the foundations lesson sets them out, and the plan arithmetic the frameworks lesson works through.

Step one: read the panel number for exactly what it is. Nielsen reports national ratings off a people-meter sample of roughly 40,000 US homes, and Nielsen sells the currency it measures, so the vendor and the yardstick are the same organisation. At national level, on a high-audience program, that sample behaves well. One layer down it stops behaving. Local market panels run in the hundreds to low thousands of homes, so a low rating in a mid-size DMA can rest on a dozen or so reporting households, and two dayparts you are trying to rank against each other often have confidence intervals that overlap completely. For Old Spice the national delivery was solid. Any market-by-market claim about where the spot "worked" was noise dressed as a decimal. And in 2010 the panel counted in-home linear viewing only, while the campaign's second life, tens of millions of YouTube views in its first week and the 186 response films in July that pushed the brand's channel to the most-viewed on the platform at the time, sat entirely outside the currency. Treat the panel figure as a floor on exposure, not a total.

Step two: build the response side off airtime logs, not off a monthly report. Stations issue as-run affidavits giving the second each spot cleared. Join those timestamps to minute-level site sessions and brand query volume, model a baseline for that minute of that weekday, and count the excess. Two things decide the answer before any analysis happens. First, the window. Published spot-level work generally finds brand search response peaking inside a minute or two and largely decayed within ten, so a three-minute window loses the viewer who searches at the next commercial break, while a thirty-minute window imports every other cause on earth. Pick the window in advance and keep it fixed across the flight, or you are choosing your own result. Second, the baseline. A 2am cable spot produces enormous percentage lift on trivial absolute volume, which is how weak dayparts get promoted in agency decks. Super Bowl weekend is the mirror image: baseline traffic is abnormal in both directions, so the counterfactual is the hard part and the arithmetic is trivial.

The join itself is a plumbing problem more often than an analytics one. You need one clock and one identity across web sessions, search referrals, app opens and orders. Segment (a customer data platform vendor, so read their material accordingly) built a business on that problem because the alternative is four CSV exports reconciled by hand every Monday, which no team sustains past week three of a flight.

Step three: name the contaminants before anyone quotes a number. Old Spice is a useful case precisely because it was contaminated. The body wash line was being relaunched with distribution and in-store display behind it, a buy-one-get-one coupon was in market, and the July response campaign landed on top of the February broadcast weight. P&G's widely quoted figures, sales up 27% across six months and up 107% in the month following the response films, circulated as claims rather than as audited reads, and trade analysts at the time queried which base period the 107% was measured against, while scanner-tracked data over parts of the same stretch looked softer than the headline. None of that says the campaign failed; the cultural result was real and the Cannes Grand Prix followed. It says an uncontaminated causal number was never available, and anyone who quoted one was asserting more than the evidence carried.

There is a second-order effect that quietly costs broadcast its budget. A viewer who sees a spot and types the brand into a browser bar arrives as direct traffic. Your analytics stack files that under direct or organic, and the channel report then shows TV contributing almost nothing while "organic" grows. The measurement artifact, not the media, is what loses the argument in the next planning meeting. Spot-level response modelling and marketing mix modelling both exist to correct that, one at minute resolution, one at weekly resolution across the whole plan.

Marketing Mix Modeling Explained

Watch on YouTube

Step four: write the verdict as a range with a date and a test attached. A defensible reconciliation on a campaign like this reads: broadcast weight in weeks one to three drove an incremental response of between X and Y, measured in a fixed ten-minute post-spot window, with a coupon drop and a distribution expansion running concurrently and not separable. Then it adds a falsifiable prediction for the next flight. The cheapest way to get that is a geographic holdout: withhold TV weight from a set of matched markets, run everything else identically, and compare. Ten to fifteen percent of your GRP weight held back buys you a clean read that no amount of post-hoc modelling will produce.

CMO action items:

  • Require as-run affidavit data to be delivered to your analytics team, not just to accounts payable for billing verification. Without exact air times there is no response side of the ledger at all.
  • Fix the response window and the baseline method before the campaign launches, and put both in the measurement plan the agency signs. Post-hoc window selection is the most common way a mediocre flight is reported as a strong one.
  • Build a geo holdout into any flight above roughly $5 million. The forgone reach in those markets is the price of knowing what the other markets actually did.
  • Ask which contaminants were in market during the read. If nobody on the call can list the promotions, price changes and distribution moves running alongside the spots, the number on the slide is not a verdict.

Common mistakes that kill results:

  • Comparing local market panel ratings as if they were precise. At DMA level, small differences sit inside sampling error, and rebuilding a plan around them is expensive superstition.
  • Quoting a lift percentage without the absolute base underneath it. Doubling 400 visits and adding 40,000 visits look identical in a percentage column and are not the same event.
  • Letting the digital dashboard adjudicate broadcast. Direct and branded-search traffic is where TV response goes to be misfiled, and a channel report built on last-click will always show the answer its attribution model was designed to produce.
  • Reporting a single causal number when the window was contaminated. Old Spice's 107% is the standing example: a real result, a real campaign, and a figure that could never bear the weight it was asked to carry.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Commission independent MMM built on three-plus years of data, validated by incrementality tests
See the full action playbook →