+100 XP

Cookieless & data clean rooms: foundations & core concepts

Third-party cookies did not die on schedule, and that is the most misread fact in marketing technology right now. Google reversed its plan to remove them from Chrome in July 2024, and in April 2025 confirmed it would not ship a standalone prompt asking Chrome users to opt out. Cookies still work in Chrome, which carries roughly two thirds of global browser usage.

Everything around them moved anyway. Safari has blocked third-party cookies by default since 2020, Firefox since 2019, and Apple's App Tracking Transparency, live with iOS 14.5 in April 2021, put cross-app tracking behind a permission prompt that most users decline. Meta told investors in February 2022 that ATT would cost it around $10 billion of revenue that year. Consent rules in Europe keep narrowing what you may collect before someone clicks accept.

The infrastructure the industry built in response is the data clean room, and it is the object this module is about. This lesson defines it: what it is, what it answers, what it will never answer, and how two companies match customers when neither will hand over a customer list.

What are we actually talking about?

A third-party cookie is a small file dropped onto a browser by a domain other than the site being visited. Advertisers used them for two decades to follow people across sites, build behavioural profiles and retarget. Look at a pair of running shoes on a retailer's site, see those shoes on a news site an hour later: that was a third-party cookie.

Cookie deprecation is browsers refusing to store or send those files. Two of the three major browsers already do. So the honest picture is a split one: cookies function in Chrome, they do not function in Safari or Firefox, mobile app tracking is broken independently of browsers, and the replacement plumbing has already been built and paid for. Google's reversal changed the deadline, not the direction.

A data clean room is a controlled environment where two or more parties combine first-party data under rules agreed in advance, and where the only thing allowed out is aggregated results. Each side contributes records it already holds lawfully. The matching and the arithmetic happen inside. Neither side can read, export or re-identify the other's rows.

Three properties separate it from a shared database. Raw records never cross: what one party contributes stays invisible to the other. Queries are restricted to an approved list of operations and outputs, agreed before anyone runs anything. And results below a minimum group size are suppressed, usually a floor in the tens of users, so you cannot keep narrowing a query until it describes one household.

That makes a clean room a measurement and audience environment, not a data transfer. If what you want is to own the other party's data, this is the wrong purchase.

Four sub-concepts every CMO must internalize

1. first-party data is now competitive infrastructure

First-party data is what you collect directly from customers with their consent: purchase history, loyalty behaviour, logged-in browsing, email engagement. You hold it, and it does not evaporate when a browser vendor changes a default.

Apple is the clearest demonstration of why this matters. It restricted third-party tracking across iOS while running its own advertising business on signals it owns outright: App Store search queries, download history, account-level data. The company that tightened the rules was also the company least exposed to them, because its data came from its own relationship with the user. That asymmetry is the whole argument for building an authenticated customer base.

2. identity resolution: the glue layer

Without a cookie stitching sessions together, you need another way to recognise the same person twice. Identity resolution matches signals such as hashed email addresses, phone numbers or device identifiers into one persistent profile. The hashed email has become the de facto key, and the two widely deployed standards built on it are The Trade Desk's Unified ID 2.0 and LiveRamp's RampID.

The friction is worth understanding before you assume email solves everything. Apple's Hide My Email issues a unique relay address to each service a customer signs into, so the address in your CRM matches nothing in anyone else's file. Every relay address, every typo, every work-versus-personal split is a customer who exists in both datasets and matches in neither.

3. data clean rooms: the mechanics

There are three shapes on the market. A platform-owned clean room sits inside a walled garden and runs on that platform's terms and its data. A neutral cloud clean room runs on infrastructure neither party owns: Snowflake, a vendor in this market, added clean room capability through its acquisition of Samooha in January 2024, so customers already storing data there can open a room with a partner without moving anything. A decentralised clean room, the model InfoSum sells, never pools the data at all: each party keeps its records in its own environment and only encrypted, non-identifying keys cross the boundary to produce the match.

The matching itself is less mysterious than the marketing suggests. Emails and phone numbers are normalised and hashed, usually with SHA-256 and a shared salt, so the same address produces the same unreadable string on both sides. Deterministic matching pairs identical hashes. Probabilistic matching infers a link from weaker signals such as IP address and device pattern, and it is a statistical guess rather than a fact. Cryptographic methods including private set intersection let two parties compute the size and composition of their overlap without either learning who sits outside it. Aggregation floors and, in some systems, deliberate statistical noise stop anyone reverse-engineering an individual out of a series of narrow queries.

One caveat that lawyers care about more than engineers do: hashing is pseudonymisation, not anonymisation. The same email always yields the same hash, which is precisely why matching works, and which is why regulators in the EU and UK generally treat hashed identifiers as personal data. The contract governing a clean room does as much work as the cryptography.

4. contextual targeting as a parallel strategy

Contextual targeting places ads against the content of the page rather than the history of the reader. Someone reading a trail running review sees trail shoes. No user data required, no identity to resolve, no consent dependency.

This is the oldest targeting method there is, getting a second look because the newer method got harder. Modern versions read page meaning through language models rather than keyword lists, which handles brand safety better than the crude blocklists of a decade ago. What contextual cannot do is tell you whether the reader bought anything. That question belongs to the clean room, which is why the two sit side by side rather than competing.

How Data Clean Rooms Work

Watch on YouTube

What a clean room can and cannot answer

what comes out

Overlap questions, first: how many of your customers also buy from this partner, how much they spend there, how the overlap group differs from the rest of your base. Deduplicated reach and frequency across two datasets that each only see their own half. Exposure-to-outcome reporting within the matched set, the closest most brands get to closing the loop between a media impression and a purchase they did not observe. And, where the agreement permits activation rather than measurement alone, audience outputs: suppression lists, seed segments, groups pushed back into a buying platform without a customer list ever changing hands.

what never comes out

Row-level records. Anything about the people outside the overlap, who are usually the majority. Causality, unless someone designed a holdout before the campaign ran, because a clean room correlates exposure with purchase and does not manufacture a counterfactual. Small slices, which the aggregation floor suppresses on purpose. And nothing at all about a competitor, since the only data in the room is data both parties agreed to contribute.

where the match rate decides everything

Every number a clean room produces describes the matched population, not your customer base. Match rates are rarely close to complete; a brand file meeting a retailer file often overlaps by a fraction, and that fraction skews towards loyalty enrollees and heavy buyers, the people most willing to identify themselves. Read a clean room result without asking for the match rate and the denominator, and you are reading a statement about your best customers and mistaking it for a statement about your market.

Knowledge check

1. What best describes the core function of a data clean room?

2. Why is first-party data described as 'competitive infrastructure' rather than just another data source?

3. A third-party cookie differs from first-party data primarily because:

MULTIPLE CHOICE

4. Select ALL statements that are accurate about cookie deprecation and its impact.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL characteristics that make a data clean room 'privacy-safe.'

Select all the correct answers.

The Future of Digital Advertising Without Cookies

Watch on YouTube

CMO action items

  • Audit your first-party data before any clean room conversation. If you cannot state how many authenticated, consent-verified customer records you hold and which identifiers each record carries, you cannot evaluate a partner's proposal, because you cannot predict your own match rate.
  • Do not read Google's reversal as permission to slow down. Safari and Firefox blind spots persist, ATT already broke mobile tracking, and consent rules keep tightening. Teams that paused post-cookie work in 2024 have lost two years of authenticated audience building they cannot buy back.
  • Make match rate and aggregation threshold standing questions for every partner and vendor. Both numbers determine what any future analysis can say, and both are knowable before you sign.
  • Learn the object by using it once. A single small test connecting your CRM to one partner's data will teach your team more about identifiers, legal review and data readiness than any vendor deck.

Common mistakes that kill results

  • Treating the clean room as a technology purchase rather than a data partnership. The software is largely commoditised. The value sits in whose data you match against. Buying a Snowflake or InfoSum deployment before you know which partner's overlap you actually want is paying for an empty vault.
  • Confusing hashed with anonymous. Hashing hides the address from a human reader; it does not remove the record from the scope of privacy law. Legal review of a first clean room agreement routinely runs longer than the technical integration, so start the data processing agreement and the consent audit at the same time as the connector, not after it.
  • Using a clean room only to explain the past. Most brands stop at reporting on finished campaigns. The output is more useful as an input: suppression of people you already converted, seeds for prospecting, budget shifts based on observed overlap.

Key Takeaways

  • Cookies survived in Chrome after Google's 2024 and 2025 reversals, but Safari, Firefox and Apple's ATT removed a large share of the signal permanently, so the shift away from third-party identifiers continues on its own momentum.
  • A data clean room is a controlled environment where two parties combine first-party data under pre-agreed rules and only aggregated output leaves: no raw records out, restricted queries, minimum group sizes.
  • Matching without a shared identifier runs on normalised hashed emails and phone numbers, deterministic where hashes align and probabilistic where they do not, with private set intersection and aggregation floors protecting the non-overlapping records.
  • Clean rooms answer overlap, deduplicated reach and exposure-to-outcome questions inside the matched set. They do not deliver row-level data, causality without a designed holdout, or anything about people outside the match.
  • The match rate is the number that governs every other number. Ask for it, and for the denominator, before you act on a clean room result.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Run a clean room pilot with a retail partner within 90 days
  • Start legal and consent review in parallel with clean room technical setup
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.