+100 XP

CDP & first-party data: frameworks & methodology

Before you sit through a single vendor demo, you owe yourself three documents: the identity resolution ruleset, the event schema, and the consent model. Write them and the demos become short, because you are asking vendors to execute a spec you already own. Skip them and you will buy someone else's data model, discover in month seven that it merges a household into one person, and pay to unpick it. This lesson is the method for writing those three documents, plus the arithmetic that tells you how fast your profiles actually need to update.

Assume the persistent profile the foundations lesson describes. The question here is how it gets built, from which keys, under which permissions.


Framework 1: the identity resolution ruleset

Identity resolution is a ranked list of match keys plus the rules for what happens when two of them disagree. Write it as a table before anyone configures anything.

Deterministic matching joins on a hard identifier: hashed email, phone, loyalty number, account ID. Same value, same person. High confidence, and it only covers customers who have identified themselves.

Probabilistic matching infers identity from device, IP, timing and behavioural patterns. It extends reach into anonymous traffic and costs you certainty. A probabilistic model tuned for coverage will happily fuse two people who share a laptop and a home network.

The practical rule: run deterministic as the spine, add probabilistic only where you can measure the error rate against a holdout of known customers. Matching 60% of your customers correctly beats matching 95% of them with a model you cannot audit, because you never see the bad matches. They show up as a customer receiving another person's order confirmation.

What the ruleset must specify:

  • Key priority. Which identifier wins when the loyalty record says one thing and the web session says another. Usually the one captured most recently under authentication.
  • Normalisation before hashing. Lowercase, trim whitespace, strip plus-addressing if you decide to, and apply the same rules everywhere. A hash of " Sarah@Example.com " matches nothing. This single detail breaks more onboarding pushes than any modelling choice.
  • Identifier caps per profile. Segment's profile unification (Segment sells the CDP under discussion) lets you limit how many values of each identifier type a profile can hold. This is your guardrail against the monster profile: a shared in-store iPad or an office IP that accumulates thousands of anonymous IDs until one record contains half your traffic.
  • An unmerge path. Merging is trivial; separating two people who were fused six months ago is only possible if you kept the raw events immutable and can re-derive the graph. Never let the CDP overwrite source events with the merged result. This is the single most expensive thing to retrofit.
  • Guest checkout handling. Retailers routinely see a large share of orders placed without an account. Decide whether an order email creates a profile, and whether that profile can be marketed to before consent exists (usually no).

Two edge cases worth naming in the document: the shared household email, common in grocery and family accounts, where one profile hides two very different shoppers; and B2B shared inboxes (orders@, accounts@), which should be excluded from person-level resolution entirely and mapped to an account instead.

When you need identity to reach beyond your own properties, that is a different tool. LiveRamp, which sells identity resolution and data collaboration, resolves offline PII to a pseudonymous RampID so you can match your customers against a partner's or a publisher's without exchanging emails. Useful for measurement and clean-room work, and not a substitute for your internal graph.

What is a Customer Data Platform (CDP)?

Watch on YouTube

Framework 2: the schema you specify

Every CDP wants the same shapes: who the person is, what they did, where they did it, and which account they belong to. Segment's spec formalises this as identify, track, page and screen, and group calls, and most competing platforms map onto the same structure. Your job is to fill it in with your business, in writing, before implementation.

Rules that hold up:

  • Name events object then action, past tense: Order Completed, Product Added, Subscription Cancelled. Mixed conventions produce three events for the same behaviour and segments that quietly under-count.
  • Cap the list. Most retail or subscription businesses need somewhere between 25 and 50 events, not the 300 an eager analytics team will propose. Every extra event is a thing to instrument, test, version and pay for.
  • Store raw facts as events, derive traits downstream. Keep purchases as purchases; compute lifetime value, propensity and tier from them. Teams who write computed segments straight into profile traits lose the ability to recompute history when the formula changes, and the formula always changes.
  • Version the schema and block violations at the gate. A tracking plan that rejects malformed events is worth more than a dashboard that reports them after the fact.
  • Reserve fields on every event for consent purpose, source system and capture timestamp. Retrofitting provenance onto two years of events is not realistic.

The failure mode: schema drift. A mobile release renames one property, nobody notices for five weeks, and the abandoned-cart audience silently drains. Put a weekly volume check on your top ten events and alert on deviation.


Framework 3: the consent model

Consent is not a boolean on the profile. It is a set of purposes, each with its own state, and it belongs in the schema from day one. GDPR enforcement has now produced well over five billion euros in cumulative fines since 2018, and the operational risk is closer to home: a suppression that fails to propagate.

Specify at least these purposes separately: transactional messaging, marketing email and SMS, on-site personalisation, and sharing with advertising platforms. A customer can accept the first two and refuse the last, and your activation layer has to respect that per destination.

For each consent record, store the purpose, the timestamp, the capture point, the jurisdiction, and the version of the notice text the customer saw. Without the notice version you cannot prove what was agreed.

Then define propagation. When someone withdraws, what is the maximum time before the suppression reaches your email tool, your personalisation engine, and your ad platforms? Put a number on it, in hours, and test it quarterly with a seeded record.

The edge case that catches people: audiences already pushed outbound. Once hashed emails sit in a walled garden custom audience or with an onboarder, they refresh on that platform's cadence, not yours. Your levers are a suppression list, a deletion instruction, and a contract clause that obliges the partner to act on it. Write the clause before you sign, not after the complaint.

Jurisdiction adds a second dimension. The EU wants consent before non-essential identifiers are set; several US states run on opt-out, with Global Privacy Control signals to honour. The same person can carry different flags in different regions, and your ruleset should say which applies when a profile has addresses in both.

Knowledge check

1. According to the lesson, what fundamentally distinguishes a CDP from a CRM?

2. Why does the lesson argue that having only data storage does not constitute a real CDP strategy?

3. What is the core methodological principle behind collecting first-party data, as described in the lesson?

MULTIPLE CHOICE

4. Select ALL statements that correctly describe the layers of a CDP as defined in the lesson.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL examples that would count as first-party data according to the lesson.

Select all the correct answers.


The latency budget

Work backwards from the use case instead of buying real-time everywhere.

Take cart abandonment. The behavioural window is short, so the chain of event capture, profile update, segment membership, and message send has to complete inside roughly 30 to 90 minutes. Now measure your actual chain. If profile recomputation runs hourly and the email platform syncs every four hours, you are outside the window and no amount of streaming ingestion fixes it, because the bottleneck is downstream.

Compare that with a win-back campaign for customers dormant 120 days. A nightly batch is fine. Paying streaming prices for it is waste.

Second-order consequence: real time multiplies cost and failure surface. CDPs price on monthly tracked users or event volume, so pushing every interaction through the streaming path inflates both the bill and the number of places a bad deploy can corrupt profiles. Pick the two or three journeys that genuinely need sub-minute updates and batch the rest.

One ordering trap: if a merge lands after the message fires, the customer gets the anonymous-visitor version of an email you meant to personalise. Decide whether your triggers wait for identity resolution or run on the pre-merge profile.


Real-world cases

Nike shows what the ruleset buys you. Nike Direct was about a quarter of revenue in FY2017; by FY2022 it reached $18.7 billion, roughly 44% of the total. The mechanism relevant here is SNKRS launch access: deciding which members see a limited release requires a profile that joins app behaviour, training activity and purchase history to one authenticated person, with enough confidence to refuse everyone else. That decision is only as good as the match rules underneath it. Nike has since rebuilt wholesale relationships, and the first-party data engine stayed.

How Nike Uses Data to Drive Business

Watch on YouTube

Starbucks shows the schema side. Rewards runs past 30 million active US members, and the profile joins mobile orders, in-store card taps, app engagement and contextual signals like time of day and weather. What makes personalised offers work is not the volume: it is that a store transaction and an app order resolve to the same member through a deterministic loyalty identifier at the point of sale. Take that identifier away and the same data becomes two disconnected halves.


CMO action items

  • Write the identity ruleset as a one-page table (keys, priority, confidence threshold, unmerge policy) and make each vendor mark up your table in the demo. The ones who cannot are telling you something.
  • Pick five audience definitions with direct revenue linkage, then trace each one back to the specific events and identifiers it requires. Any segment that depends on data you do not collect is a collection project, not a CDP project.
  • Set a data quality SLA with engineering: match rate against known customers, acceptable profile staleness per journey, and what happens to an event that fails identity resolution. Numbers in a contract, not a slide.
  • Seed a test record and run a full deletion request end to end, before go-live and every quarter after.

Common mistakes that kill results

Inheriting the vendor's data model. Default match rules are tuned for demo coverage, not for your household structure or your guest checkout rate. If you do not hand over a spec, you adopt theirs by silence.

Merging without an unmerge path. Overwriting source events with the merged profile feels tidy and makes every future correction a manual data project. Keep raw events immutable and treat the graph as derived.

Consent bolted on afterwards. Once profiles exist without purpose flags and notice versions, you are choosing between a backfill you cannot evidence and a suppression list that grows forever.

Ignoring decay. Email lists go stale at something like a fifth of records a year, and intent signals age in days in some categories. Budget for hygiene, re-permission and refresh, or your personalisation will speak to who the customer was 18 months ago.


Key takeaways

  • Bring three written artefacts to the vendor conversation: identity ruleset, event schema, consent model. The demo is a test of your spec, not their slides.
  • Deterministic first, probabilistic only where you can measure error against a known-customer holdout. Silent false merges cost more than missed matches.
  • Cap identifiers per profile and keep source events immutable, so a bad merge stays reversible.
  • Name events object-action past tense, keep the list to a few dozen, and derive traits from raw facts so you can recompute when the formula changes.
  • Model consent as purposes with timestamps, jurisdiction and notice version, and put an hours-level SLA on suppression reaching every downstream destination.
  • Set the latency budget per journey. Cart abandonment needs minutes; a 120-day win-back does not, and streaming everything inflates the bill and the failure surface.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Define five specific activated CDP use cases before signing any contract
  • Have marketing, not IT, own CDP data model, segments, and activation triggers
  • Budget ongoing data hygiene, re-permission, and behavioral refresh as recurring discipline
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.