+150 XP

Engagement metrics that predict churn before it happens

A subscriber who skipped a third more tracks than her own weekly average, opened the app three fewer times and saved nothing new is not a blip. She is a cancellation two or three weeks out, and a team watching individual trend lines already knows it. Spotify has said publicly that skip behaviour feeds its recommendation systems; read as a trend rather than as a taste signal, that same data is one of the earliest churn flags a subscription business has.

What follows is narrow on purpose: not whether someone converted, which is settled upstream, but whether a paying subscriber's behaviour is decaying, how fast, and how much warning that decay actually buys you.

Why churn is a lagging metric

Churn rate (the percentage of subscribers who cancel in a given period) tells you what already happened. By the time it moves, the customer and the money spent acquiring them are both gone.

So teams build leading indicators: metrics that move before the cancellation event. For subscription media they cluster into four families:

  1. Frequency (are they showing up?)
  2. Recency (how long since the last real interaction?)
  3. Breadth and depth (how much of the catalogue, and how far into each thing?)
  4. Friction (are they struggling, or drifting?)

The core engagement metrics

Session frequency

How many times a user opens the app or starts content in a rolling window, usually 7 or 28 days.

  • DAU/MAU ratio (daily actives divided by monthly actives) is the standard stickiness measure. A ratio near 50% means the average monthly user shows up roughly every other day.
  • Benchmark estimate: large consumer media apps run DAU/MAU around 25 to 35% based on recent public commentary; below 15% in a daily-use category like music is a warning sign (varies with how each platform defines an active user).

Worked example: 10 million MAU and 3 million DAU gives DAU/MAU = 30%. If one cohort drifts from 30% to 22% over two months, its churn risk has risen even though nobody has cancelled yet.

Recency

Days since the last meaningful action, not merely the last app open. This single number often out-predicts the composite scores built on top of it. A user with nine active days out of the last ten and zero in the last four is a sharper flag than one with a gently sagging monthly total.

The Financial Times built its engagement measure, RFV (recency, frequency, volume), on this logic, and uses the score to anticipate who will renew rather than to describe traffic. Volume there is articles read; at Sky it would be hours viewed; at Spotify, listening time. The frame travels, the units do not.

Skip rate

The share of content units abandoned before a completion threshold, for example a track skipped inside 30 seconds.

Skip rate = (tracks skipped early) / (total tracks played)

A rising trend points to weakening fit between the recommendations and that user's taste. Absolute levels barely compare across platforms because thresholds differ, and for some listeners heavy skipping is just how they browse. The trend against their own baseline is the signal.

Completion rate

For video, the share of an episode or film watched through; for podcasts, the share of the episode consumed. Sky's mix of live sport, box sets and linear channels makes a single blended completion figure useless: 40% of a three hour match broadcast and 40% of a 50 minute drama describe two different behaviours. Cut it by content type or do not report it.

Content breadth

Distinct artists, shows, sections or genres per period. A narrowing repertoire (the same five tracks, or an FT reader who only ever opens one section) often precedes boredom-driven cancellation. Breadth has a second-order effect too: broad users survive a weak content month because they have somewhere else to go inside the product, so a catalogue gap hits your narrow users first.

Habit signals

Playlist creation, saves, follows, push open rate. Someone who stops saving has left the habit loop while the login count still looks fine.

Building a leading-indicator churn model

Teams combine these signals into a churn propensity score, a per-user probability refreshed weekly.

Simplified logic (illustrative, not any specific company's model):

risk_score = 
    0.3 * normalized(decline_in_session_frequency) +
    0.25 * normalized(increase_in_skip_rate) +
    0.2 * normalized(decline_in_session_length) +
    0.15 * normalized(decline_in_content_breadth) +
    0.1 * normalized(decline_in_social_actions)

if risk_score > threshold:
    flag_user_for_retention_campaign()

Every decline is measured against that user's own history, not the population mean. A power user falling to average looks nothing like a light user staying flat, which is why cohort averages miss the early weeks.

Two structural limits are worth knowing before you trust the output.

The first is cold start. A relative-decline model needs six to eight weeks of individual history, and the heaviest cancellation period in most subscription businesses is the first billing cycle. For new subscribers you have to score activation milestones instead: three sessions in week one, a first saved item, a first return visit without a push.

The second is the counter-example the model scores as safe. A Sky Sports subscriber who watches every fixture of a season and cancels the week after the final had peak engagement right up to the exit. An FT reader who took an annual subscription for one transaction behaves the same way. Goal-completion churn needs event and contract boundaries in the score, not decay curves.

What engagement decay costs, and what the model cannot see

Shorter expected lifespans compress the value side of the equation, which the lifetime value lesson turns into money against the acquisition costs the CAC lesson builds. The question this lesson answers is different: how many weeks of warning you get, and how reliable the warning is.

Two things sit outside the reach of any engagement score. Involuntary churn (expired cards, failed payments) can account for a substantial slice of gross cancellations in subscription businesses, and the users it takes are often fully engaged; that problem belongs to billing and dunning, not marketing. Price-shock churn is the other: a subscriber can be a daily user and still leave over a tariff increase, and no behavioural signal will show it in advance.

You also cannot claim a save without a control. Hold back 5 to 10% of flagged users untreated and measure the difference in survival at 30 and 90 days. Without that holdout, every retention campaign appears to work, because flagged users regress towards their own mean anyway.

For deeper background on subscription funnel economics, see this overview from NYU Stern's coverage of subscription business metrics or industry primers from a16z on SaaS and subscription metrics, whose frameworks map directly onto media subscriptions.

Knowledge check

1. Why is churn rate considered a lagging metric rather than a useful early warning signal?

2. A subscriber plays music less often, opens the app fewer times, and stops saving playlists. Which category of leading indicator does this combination primarily represent?

3. What does the DAU/MAU ratio primarily measure, and why is it useful for churn prediction?

MULTIPLE CHOICE

4. Select ALL correct answers about the three families of leading indicators for subscriber churn described in the lesson.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about why media companies build leading indicators instead of relying solely on churn rate.

Select all the correct answers.

What "good" looks like: sector benchmarks (estimates)

MetricHealthy range (estimate)Warning zone
Monthly churn, music/video streaming2 to 4% (as of recent industry commentary)above 6 to 7%
DAU/MAU, daily-use category (music)25 to 35%below 15%
Skip rate, music streamingvaries by platform; the trend beats the levelsustained increase over 4+ weeks
Share of paying base with no meaningful action in 14 dayslow single digits in a daily-use productdouble digits, or rising month on month

These are directional estimates for context, not disclosed figures. Check each company's investor reporting (Spotify's quarterly shareholder letters, for instance) for actual numbers, and read the definitions before comparing.

🎬 [VIDEO: "How Spotify Uses Data to Personalize Your Music" - youtube.com - a walkthrough of how listening behavior feeds Spotify's recommendation and retention systems]

The marketing action layer

Detecting risk is half the job. The useful tiering is by lead time and contract shape, not by discount size; the offer ladder itself belongs to the monetisation and churn-reduction lesson.

  • Early drift, four or more weeks out: change what the user sees. Surfacing new releases, a personalised playlist, an editorial newsletter matched to the sections they actually read. It costs nothing and it does not remind anyone that cancelling is an option.
  • Mid-stage: targeted email and push, tested for tone and for send timing against the holdout.
  • Late stage, near a renewal or contract boundary: hand to retention, with the score attached so the team knows whether they are saving a lapsed habit or a completed goal.

Contract shape decides how much of this is even actionable. On monthly plans you can intervene continuously. On Sky's minimum-term TV contracts or an FT annual subscription, cancellations bunch at the term boundary, so a score that flags decay in month four is only useful if it survives in the record until the renewal window opens. Run the model against renewal dates, not against calendar weeks.

One failure mode to plan for: waking a dormant subscriber with a "we miss you" message can prompt the cancellation it was meant to prevent. For deeply inactive users, quiet re-engagement through content beats direct address.

Key takeaways

  • Churn is a lagging metric; frequency, recency and breadth move weeks earlier and are what marketing should monitor.
  • Recency of a meaningful action, not just app opens, is often the strongest single predictor. The FT's RFV score is built around it.
  • Score each user against their own baseline, and accept that the model is blind for the first six to eight weeks, exactly when churn is highest.
  • Goal-completion churn (the season ends, the deal closes) shows no decay at all, so build contract and event boundaries into the score.
  • Involuntary and price-driven cancellations sit outside engagement data entirely; without a holdout you cannot tell a real save from natural regression.