Reading engagement metrics that predict retention before it happens
parkrun hands every registrant a barcode. Scan it at the finish line and the attendance is stamped with a date, a location and an ID that persists for life. One design decision, made once, means the organisation can watch a participant drift away weeks before that person would describe themselves as having stopped: one missed Saturday, then three, then a gap that quietly never closes. Most public programmes could see the same drift. They simply never instrumented the moment it happens, so the first hard evidence of lapse is the renewal that does not arrive.
Leading indicators move before the outcome you care about. Lagging ones confirm it afterwards. This lesson is about which signals reliably lead, how far ahead they lead, and what you have to log at sign-up to have them at all.
Why retention needs leading indicators
Annual retention is arithmetic on the past. It is fine for a board pack and useless for intervention, because the members it counts as lost left months ago. (How your figure compares with anyone else's is the benchmarking lesson's problem, and the peer set it insists on matters more than the number.)
Wikimedia publishes monthly counts of active editors, defined as registered accounts making five or more edits in a month. That is a clean, honest measure, and it still resolves at a month. The behaviour that decides whether a new editor is still there in a year happens in their first session, in the first fifteen minutes.
The three leading indicators
1. The frequency gap, not the attendance rate
For any recurring activity, the useful number is weeks since last participation, held per person, updated automatically. An attendance rate averaged across a cohort hides the thing you need: individuals whose personal interval has stretched.
parkrun's structure makes this readable. Events run weekly, at the same time, in thousands of locations across more than twenty countries, and the barcode ties every finish to the same record. Two behaviours are worth separating in that data. First, registrants who never produce a first scan at all: a large share of sign-ups, and a different problem from lapse, because nothing has yet been built to lose. Second, established participants whose interval doubles. Someone averaging a run every ten days who goes to thirty is not on holiday.
parkrun also logs volunteer occasions against the same barcode, with a milestone at twenty-five. Instrument that crossover deliberately. A participant who has taken a marshal shift has a relationship with the event that survives an injury; a participant who has not, does not.
2. The outcome of the first contribution
Wikimedia's editor decline after 2007 is the best-documented case in the sector, and the research (Halfaker and colleagues, 2013) points at newcomer experience rather than newcomer supply: good-faith first edits reverted by automated tools and tightened patrolling, at scale. The predictive signal is not how many edits a newcomer makes. It is whether their first edits survive, and whether anyone spoke to them like a person. The Teahouse, opened in 2012 as a help space for new accounts, exists because that signal was found and acted on.
Generalise it. Log the result of a first attempt, not just the fact of one: was the form rejected, the application returned, the submission silently binned? Most organisations record the count and throw away the outcome, which is the half that predicts anything.
3. Depth per session against time spent
Khan Academy's guidance to teachers has centred on roughly thirty minutes a week per student of mastery-based practice, with reported gains associated with hitting it. That gives a threshold you can instrument: not "logged in", but "crossed the dose that correlates with progress".
The counter-example matters as much. A learner whose minutes climb while mastery points stay flat is grinding material already known, or stuck. Rising time with static progress is a lapse signal wearing engagement's clothes, and a dashboard that reports minutes alone will read it upside down.
Worked example: the twelve-month dataset
A programme with 8,000 registrants tracks a cohort of 500 monthly across three logged behaviours.
| Month | Contact opens | Session attendance (% of cohort) | Avg. volunteer occasions/member |
|---|---|---|---|
| Jan | 32% | 18% | 2.1 |
| Apr | 29% | 15% | 1.8 |
| Jul | 24% | 11% | 1.2 |
| Oct | 19% | 7% | 0.6 |
By month ten all three decay together. Expressed as a ratio against the same period a year earlier, volunteer occasions run at 0.6 / 2.1, roughly 0.29. Halved year on year is the point at which most programmes that track this see non-renewal follow within a quarter.
Wait for the renewal cycle and you learn this in month twelve. Score it from months four to seven and you have a five to eight month window.
A simple composite risk score
def lapse_risk_score(open_rate, event_pct, vol_index, weights=(0.3, 0.4, 0.3)):
# normalize each input to 0-1 scale against your own historical baseline
# lower score = higher risk
return (weights[0] * open_rate) + (weights[1] * event_pct) + (weights[2] * vol_index)
# example: month 10 cohort
score = lapse_risk_score(open_rate=0.19, event_pct=0.07, vol_index=0.29)
print(round(score, 3)) # 0.190Weight the inputs by what actually precedes lapse in *your own* history. Advocacy bodies often see attendance go first; credentialing bodies often see committee hours go first, because service is tied to certification cycles.
Knowledge check
1. What is the fundamental distinction between a leading indicator and a lagging indicator in the context of membership retention?
2. Why does the lesson argue that an annual retention rate of 82% is 'useless for intervention'?
3. Based on the described pattern of disengagement, why would monitoring email open rate decay be considered an earlier warning sign than tracking volunteer hours?
4. Select ALL correct answers about why organizations should track leading indicators of member disengagement.
Select all the correct answers.
5. Select ALL correct answers about the relationship between the three engagement metrics discussed and member lapse.
Select all the correct answers.
Turning signals into intervention funnels
A score with no action attached is a dashboard. MapMapUsing software to automate repetitive marketing tasks and campaigns, enabling personalisation at scale across channels like email, web, and social.View full definition → each tier onto the stages the funnelfunnelThe customer journey from awareness to purchase, typically Awareness, Interest, Consideration, Decision, Action, with prospects narrowing at each stage.View full definition → lesson already names, and fix the cost per contact before you switch it on.
Tier 1, early decay: automated re-engagement sequence. Near zero marginal cost.
Tier 2, two signals falling together: a call or message from a real person, ideally a peer rather than staff. Association and nonprofit literature commonly puts new-supporter acquisition at three to five times the cost of a save (an estimate, and it varies wildly by channel).
Tier 3, everything near zero: a direct conversation built on a reminder of what the person gets, not a discount. Discounting teaches the whole cohort to wait for one.
Is the save worth the call?
Take the value of a retained supporter as the LTVLTVLifetime Value: the total revenue (or profit) a customer generates throughout their entire relationship with your business.View full definition → lesson builds it, then do the precision arithmetic your model quietly needs. Suppose 12% of 8,000 registrants lapse in a year: 960 people. Your score flags the riskiest 20%, or 1,600, and captures half the real lapsers inside that group. You have 480 true positives and 1,120 people who were going to stay anyway. At £20 of loaded staff time per call, that is £32,000 to reachreachThe number of unique people exposed to your message in a given period. Unlike impressions, reach counts each person once, no matter how often they see it.View full definition → 480 at-risk supporters, about £67 each. Tier 2 survives that maths. A £300 home visit does not, which is why intervention cost has to be designed against the model's precision rather than its recall.
🎬 [VIDEO: "Membership Renewal Strategy: Predicting and Preventing Lapse" - youtube.com - search for association management webinars covering renewal funnel design and early-warning engagement tracking]
What to watch out for
Email opens stopped being a clean signal in September 2021, when Apple's Mail Privacy Protection began pre-fetching images for Apple Mail users and registering opens that no human performed. Even the vendors publishing open-rate figures now caveat them; Mailchimp, which sells the sending software, is one. If your risk model still weights opens at 0.3, it is weighting a machine.
Separate demand-side lapse from your own supply failure. A parkrun event cancelled for frozen ground or a volunteer shortage produces the same zero in the data as a participant who has lost interest, and running a win-back campaign at people whose event you cancelled is worse than doing nothing. Join attendance data to your own outage log before scoring anyone.
Seasonality does the same trick. Academic cycles, dark winter mornings and fiscal year-ends all produce dips that reverse unaided. Compare with the same period last year, never with last month.
Then watch what the metric does to your staff. Once a team is measured on open rates, someone starts re-sending to non-openers; once it is measured on volunteer hours, hours get logged generously. A leading indicator that people are paid to move stops leading anything. Keep the score internal to the analytics function, publish the outcome it predicts, and revalidate the weights against real lapse data every year.
Key Takeaways
- The signal is a per-person frequency gap, not a cohort average: weeks since last participation, updated automatically, flags drift months before a renewal date does.
- Log the outcome of a first contribution, not just its existence. Wikimedia's decline traces to reverted good-faith first edits, which a simple edit count would never have exposed.
- Rising time-on-task with flat progress is a warning, not engagement. Instrument a dose threshold (Khan Academy's roughly thirty minutes a week is one) rather than logins.
- Design intervention cost against the model's precision. Flag 20% of an 8,000-person base and most of those flagged were staying anyway, which rules out any save costing more than a few tens of pounds.
- Distinguish your own outages from real lapse, control for seasonality, and never let the team be targeted on the leading indicator itself.