+55 XP

Loyalty & retention: frameworks & methodology

Four cohorts are on the screen. January's signups have flattened at 38% still active in month six. February looks the same. March is still falling at month three and nobody can say where it lands. You have one lifecycle campaign's worth of budget and two sprints of engineering time this quarter. Which cohort do you spend it on?

You cannot answer that from a churn rate, the number the foundations lesson defines. A single monthly percentage blends cohorts of different ages, different channels and different intent into one figure that moves for reasons you cannot see. The method below gets you to a defensible answer in three moves: a curve you trust, a prediction that beats the base rate, and plumbing that fires the action while the customer is still reachable.

Reading the curve before you read the number

Build the cohort triangle first: rows are signup month, columns are months since signup, cells are the share of that cohort still active. Two properties matter, and they are independent. Level is how high the curve sits. Shape is whether it flattens.

A curve that approaches an asymptote at, say, 32% tells you a third of every cohort is a stable base, and the lifetime value arithmetic the foundations lesson sets out runs off that asymptote plus monthly contribution margin. A curve that never flattens has a terminal retention of zero. Every customer is rented, and no loyalty mechanic fixes that; the product does or it does not.

Write down which retention definition you are using and stop changing it. N-day retention counts users active exactly on day N. Bracketed retention counts anyone active inside a window (days 25 to 31). Rolling or unbounded retention counts anyone active on day N or any day after, which means a lapsed user who returns in month nine retroactively becomes "retained" in months three through eight. Rolling retention can only be revised upward, so recent cohorts always look worse than mature ones and your dashboard flatters the past. The same raw event log can produce double-digit differences at month three depending on which of the three you picked.

Then pick your cohort by slope, not level. The month with the steepest bend between weeks one and four has the most recoverable value, because that is where the population is still large. Before you commit, decompose the cohort by acquisition channel. A curve that dropped four points quarter on quarter is more often a mix shift (paid social went from 12% to 30% of signups) than product decay. Fix the wrong thing and you spend a quarter improving onboarding for a cohort whose problem was media buying.

Framework 1: RFM as the scoring layer

RFM sorts customers by recency, frequency and monetary value, each cut into quintiles. That is 125 cells, which nobody can operate, so collapse them into eight to twelve named segments with one owner and one action each. For subscription businesses, swap monetary for contribution margin and frequency for active days per month; gross revenue quintiles will rank your heaviest discount users at the top.

The methodological point most teams miss: RFM is descriptive. It says where someone sits, not where they are heading. The signal is the delta. A customer who moved from the top frequency quintile to the middle over eight weeks is a live retention case; a customer who has sat in the middle for two years is stable. Score RFM monthly, store the history, and alert on movement rather than on state.

Framework 2: activation, where the curve is steepest

Most of the gap between a good cohort and a bad one is set in the first days, before any loyalty mechanic can reach it. So define one activation event, the observable action that separates the two curves, and instrument it before you write a single email.

The arithmetic is worth doing by hand. Take 100 signups: 40 complete the activation event and retain at 55% on day 30, the other 60 retain at 12%. Blended day-30 retention is 29.2%. Move activation from 40% to 50% and, with both sub-rates unchanged, you land at 33.5%. Four points of retention from a funnel fix, not a discount.

The trap is that activators are self-selected. People who set up the integration were more committed on the day they signed up. The only way to know whether pushing them there causes retention is a randomised holdout on the onboarding sequence itself. Drift, which sells conversational marketing software, built its product around firing a scripted play inside the live session based on who the visitor is; the same logic applies post-signup, and the same evidence standard applies to both.

Framework 3: churn prediction that earns its compute

Start with the base rate. If monthly churn is 3%, a model that predicts "nobody churns" is 97% accurate and worth nothing. Judge on lift in the top decile: if the top 10% of scored accounts contains 35% of next month's churners, that is 3.5x lift and a target list a human team can work.

Four things break these models in practice:

  • Leakage. "Visited the cancellation page" and "downgraded plan" are the outcome wearing a feature's clothes. The model scores brilliantly in backtest and warns you 40 minutes before the customer leaves.
  • Horizon mismatch. A model tuned to predict churn three days out is useless if the save play needs a CSM call and a two-week scheduling window. Match the prediction horizon to the action lead time, then hold it fixed.
  • Wrong question. Classifiers answer "will they churn inside this window". Survival models (Kaplan-Meier for the descriptive curve, Cox for the drivers) answer "when", and handle the customers who have not churned yet without throwing them out of the sample.
  • Targeting risk instead of persuadability. The highest-risk decile is full of people already gone. Uplift modelling scores the difference between outcome with treatment and without, and it routinely finds that a chunk of your target list would have stayed anyway while another chunk churns faster when you contact them.

Survey sentiment sits badly in these models. NPS, introduced by Fred Reichheld with Bain in 2003, arrives from a self-selected minority of your base and is stale within weeks. As a feature it is sparse; behavioural features beat it almost every time.

How Starbucks Built the World's Best Loyalty Program

Watch on YouTube

Framework 4: the jobs lens, instrumented

Clayton Christensen's Jobs-to-Be-Done argues customers hire a product for a job. Spotify treats listening contexts as separate jobs: commute, workout, focus, winding down, and has built the catalogue and the recommendation surfaces (Discover Weekly, from 2015) around them.

The measurement version is simple. Count distinct contexts per user per week, then cut the retention curve by that count. If the one-context curve and the three-context curve separate, you have a lever worth building product for, and a lifecycle trigger worth writing: users who have only ever used the product for one job get prompted toward a second.

The plumbing: from event to lifecycle action

None of the above fires without event and identity infrastructure. Segment, a customer data platform (so a vendor of exactly the plumbing described here, bought by Twilio in 2020 for around $3.2bn), organises this around a small set of calls: track for events, identify for user traits, group for accounts.

The failure that quietly destroys cohort analysis is identity resolution. The same human arrives anonymously on the web, signs up on iOS, and appears again on an invoice under a work email. Without aliasing the anonymous ID to the user ID at login and stitching accounts to users, one person becomes three profiles. Denominators inflate, retention understates, and the churn model trains on ghosts.

Second failure: schema drift. An app release renames "Song Played" to "Track Played". The curve for that cohort falls off a cliff, the alerting fires, and you launch a win-back campaign at people who never left. A versioned tracking plan with named owners and a CI check on event names costs a day and prevents a quarter of bad decisions.

Third: latency. A dunning save on a failed card has a window of hours, not a nightly batch. Decide per trigger which needs streaming and which can wait, because streaming everything multiplies cost for no gain.

Fourth: deletion requests remove profile history and break cohort continuity. Keep an anonymised cohort counter so your denominators survive the erasure of the individuals inside them.

Duolingo's Growth Strategy Explained

Watch on YouTube

CMO action items

  • Publish one written retention definition (N-day, bracketed or rolling) and rebuild the cohort triangle on it. Circulate the old and new numbers side by side so nobody argues about the discontinuity later.
  • Choose this quarter's cohort by curve slope, after decomposing by acquisition channel to rule out mix shift.
  • Name one activation event, measure the day-30 retention gap either side of it, and run the onboarding push against a 10% holdout.
  • Score your churn model on top-decile lift, not accuracy, and audit the feature list for anything that only appears after the decision to leave.
  • Ask engineering for the ratio of profiles to known humans. Anything above 1.2 means your curves are wrong before the analysis starts.

Common mistakes that kill results

Reporting retention annually hides everything. The decision to leave is usually made in weeks two to eight; an annual cohort read is the autopsy, not the diagnosis.

Running save campaigns with no control group means you can never separate the intervention from regression to the mean. Every play needs a randomised holdout, and holdouts need to survive the first quarter someone asks to switch them off because "we know it works".

Treating a satisfaction score as a retention forecast is a category error. Satisfaction is a state; retention is behaviour. Build the measurement stack on events you can observe, and use survey data for diagnosis rather than for targeting.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Segment retention metrics by cohort, tier, and acquisition channel
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.