AI & ML in marketing: frameworks & methodology
Your team arrives at the planning meeting with nine AI proposals and budget for two. One is a churn model. One is a generative assistant for campaign briefs. One is a next-best-offer engine for the loyalty app. Nothing in the deck tells you which will pay back, because every deck promises the same lift. What follows is the screen you run before anyone trains anything: which decisions deserve a model, what the data has to look like first, and how to measure the result so it survives a conversation with your CFO.
Core concept: which decisions deserve a model
Prediction earns its budget only when it changes an action. Take the machinery as given (the foundations lesson covers what a model is and where prediction beats a rule) and ask something narrower than "can we predict this": what would the business do differently at 9am on Tuesday if the number existed?
Four tests, and a candidate has to pass all four:
- Frequency. The decision recurs thousands of times a quarter. A model that informs one annual budget split is a spreadsheet with extra steps and a longer procurement cycle.
- Differentiated action. You can treat customer A differently from customer B at a cost you can bear. If everyone receives the same email regardless of score, the score is decoration.
- Observable outcome, fast enough. The thing you predict is recorded unambiguously inside a window short enough to retrain against. Ninety days works. A three-year renewal cycle does not.
- Bounded error cost. A wrong product recommendation costs a few cents of margin. A wrongly withheld credit offer, or an automated message sent into a bereavement, costs considerably more, and those use cases need a human gate before they need a model.
The counter-example worth keeping in your head: enterprise deal scoring in a business that closes 200 deals a year. The prediction is technically possible, the training set is minuscule, the feedback loop runs eighteen months, and the reps override the score anyway because they were in the room. That project fails three tests out of four before anyone writes a line of code.
Key sub-concept 1: size the prize before the build
Do the arithmetic on the back of the proposal: decisions per year, multiplied by the margin at stake in each decision, multiplied by a plausible relative lift, minus the annual run cost.
Say the loyalty app sends 40 million offers a year, 3% get redeemed, and each redemption carries roughly four units of margin. That is 1.2 million redemptions. An 8% relative improvement adds about 96,000 redemptions, worth around 384,000 in margin. Now put the fully loaded cost of two data scientists, an engineer's time and the infrastructure next to it. Some use cases stop right there, which is the point.
Be honest about the lift you assume. The published 20% to 30% numbers usually come from replacing batch-and-blast with something, anything, personalised. Measured against a decent recency and frequency rule that your CRMCRMCustomer Relationship Management: software and strategy to manage and analyse customer interactions throughout their lifecycle.View full definition → team already runs, 3% to 8% relative is a more defensible planning figure. If the business case only closes at 15%, kill it in the meeting rather than in month nine.
Key sub-concept 2: the feature and data readiness test
Five checks, run before the model is approved, written up as a one-page memo by whoever owns the data pipelinedata pipelineETL (Extract, Transform, Load) is a data integration process that pulls data from sources, reshapes it into a consistent format, and writes it into a target system.View full definition →:
- Volume in the rare class. Not total rows. The count of the outcome you care about. Salesforce, which sells the scoring product, publishes floors for Einstein Lead Scoring in the region of a thousand leads and over a hundred conversions in the trailing six months. Treat a vendor's stated minimum as the point below which the product refuses to run, not the point at which it works well.
- Leakage. A feature that only exists because the outcome already happened. Cancellation-page visits in a churn model give you a beautiful offline AUC and nothing in production. The test: could this field have been populated at the moment of scoring, for a customer whose outcome you do not yet know?
- Point-in-time correctness. Training on today's profile attributes rather than their value on the date of the event teaches the model the future. Same failure, quieter.
- Coverage at serving time, not warehouse time. A field that is 90% complete in the nightly table and 40% populated in the real-time call is a different feature. Score the model on what the APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → will actually receive.
- Identity. Scores attach to the persistent profile described in the CDPCDPSoftware that unifies customer data from every source into one persistent profile that marketing, sales and service teams can act on.View full definition → lesson. If the stitching is wrong, the model is precise about the wrong person, and the error is invisible in every offline metric you will look at.
One edge case that catches teams late: a feature can be predictive and unusable. Postcode-derived affluence proxies, health inferences, anything that reconstructs a protected characteristic. Get legal into the feature list at design time, not at launch review.
Key sub-concept 3: the activation bridge
A score that lands in a dashboard has produced a report. The bridge is the automated path from model output into the system where the decision is executed: the offer layer, the bidder, the CRM queue, the app.
Starbucks built Deep Brew from 2019 under Kevin Johnson precisely as an activation layer rather than an analytics one. Offer selection for loyalty members runs through it and into the app, and the same platform reaches into store operations such as labour scheduling and inventory. The loyalty programme is what makes it possible: Rewards members account for over half of US company-operated revenue, so there is a known customer, a known moment, and a channel that can execute a per-member decision in milliseconds.
Four engineering questions the pilot usually forgets, all of which are yours to ask:
- What is the latency budget, and what does the app show if the score arrives late?
- What is the fallback when the score is missing or stale? A named default offer beats an empty module.
- Who can override, and is the override logged as training data?
- Do frequency caps sit upstream or downstream of the model? Downstream caps silently break your measurement, because the treatment group did not receive the treatment.
How Spotify's Algorithm Works
Key sub-concept 4: measuring uplift honestly
Randomise a holdout at customer level, size it before launch, and keep it for the life of the programme.
Size it with the standard approximation: subjects per arm is roughly 16 × p × (1 - p) ÷ d², where p is your baseline rate and d the absolute difference you want to detect at 80% power. On a 5% conversion baseline, detecting a 10% relative lift (0.5 percentage points) needs about 30,000 per arm. Detecting a 2% relative lift needs about 760,000 per arm. If your monthly addressable audience is 200,000, you cannot see a 2% effect, so nobody should be forecasting one.
Rank by uplift, not by propensity. The highest-scoring customers are frequently the ones who would have converted without any contact, so spending on them produces impressive response rates and no incremental revenue. Four groups exist: the sure things, the persuadables, the lost causes, and a do-not-disturb group where contact actively reduces conversion, the classic case being a win-back email that reminds a dormant subscriber that they are still paying. Read treated-minus-control within each score decile. The overall response rate of the treated group is not a result.
Two timing rules. Read at week eight or later, because novelty inflates the first month. And run champion-challenger with a pre-registered metric and a decision date, otherwise the challenger stays live until someone finds a week where it won.
Real-world cases
CASE 1: Starbucks and the offer decision
The decision Deep Brew models is narrow: which offer to put in front of an identified member at a specific moment. It passes the screen cleanly. Millions of decisions a day, per-member action through an owned channel, redemption visible within days, and the cost of a wrong offer is one wasted incentive. The second-order consequence deserves your attention more than the model does: personalised discounting teaches members to wait for the offer. Track margin per member per quarter alongside redemption rate, or you will optimise your way into a cheaper basket.
CASE 2: Salesforce Einstein scoring and the reason reps ignore scores
Salesforce sells this category, so read its constraints as product design rather than neutral advice. Einstein Lead Scoring surfaces the top factors pushing a score up or down, not just the number. That exists because the recurring failure of scoring deployments is not accuracy, it is abandonment: a bare number in a CRM field gets ignored by the person whose commission depends on their own judgement. Budget for the explanation layer and the rep training in the same line item as the model.
Machine Learning for Marketing Explained
CASE 3: the contaminated holdout, a pattern you will meet
A team launches with a clean 10% holdout. Quarter two, the CRM team runs an unrelated campaign across the whole base, including the holdout. Quarter three, someone shrinks the holdout to 2% to capture more revenue. Quarter four, finance asks what the programme delivered and the only available answer is a year-on-year comparison contaminated by pricing, seasonality and a media shift. The model may well have worked. It is now unprovable, which in a budget round is the same as not working.
CMO action items
- Run the four-test screen on every proposal in the deck, and require one sentence naming the action that changes. If nobody can write that sentence, the project does not start.
- Demand the readiness memo before releasing spend: rare-class counts, null rates at serving time, the leakage check, and confirmation that every feature exists at the moment of scoring.
- Fix the measurement design in writing pre-launch: holdout size derived from the arithmetic above, a named owner, a read date at week eight or later, and a rule that any reduction of the holdout comes to you for approval.
- Put a single person, not a team, on model degradation, with a monthly drift review and a scheduled retrain.
Common mistakes that kill results
- Optimising the model rather than the margin. Moving AUC (a technical measure of ranking quality) from 0.78 to 0.82 while pipelinepipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition → stays flat means the output is not reaching a decision, or it is reaching people who were already going to buy.
- Reporting the treated group's response rate as the result. Without the control cell, that number tells you how good the targeting was at finding buyers, not how many buyers the targeting created.
- Piloting on data the scaled version will never have. Hand-cleaned extracts, a manually stitched customer list, one analyst reconciling IDs on Fridays. The pilot works, production does not, and the diagnosis lands months later.
- Funding it as a project. Models decay as behaviour and product mix shift. Without a maintenance budget and a refresh schedule, you buy one good quarter and a slow slide that everyone attributes to the market.
Resources
- 🔗Google's Rules of Machine Learning: Best Practices for ML Engineering
Google's internal engineering guide on ML methodology, written by Martin Zinkevich, gives CMOs a concrete understanding of how production ML systems are scoped, built, and maintained so you can hold technical teams accountable.
- 🔗Harvard Business Review: How Marketers Can Use AI Without Losing the Human Touch
Raj Venkatesan and Jim Lecinski's AI Marketing Canvas provides a structured framework for mapping AI applications to specific marketing funnel stages, directly applicable to the problem-first methodology covered in this lesson.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Set model monitoring with retraining triggered when predictions drift beyond 10%
- Name a single person accountable for model performance and interpretability
Related articles
Recent articles from the blog that build on this lesson.