+100 XP

AI & ML in marketing: foundations & core concepts

Sit through three vendor demos in one week and "AI" will describe three unrelated things: a scoring model trained on your CRM, a language model writing subject lines, and a set of if-then rules with a confident deck wrapped around them. All three arrive under the same word, and only one of them learns anything from your customers. This lesson gives you the vocabulary to tell them apart: what training data is and what it silently decides on your behalf, what a propensity model produces compared with what a generative model produces, and the narrow conditions under which a prediction genuinely beats a rule you could have written yourself.


What AI and ML actually mean in a marketing context

Artificial intelligence (AI) is the broad category: software that performs tasks normally requiring human judgment. Machine learning (ML) is the subset where the system infers patterns from historical data and updates its outputs as new data arrives, with nobody rewriting the logic by hand.

A rules-based email system sends a discount to everyone who abandoned a cart, because a person decided that once. An ML system works out from past behaviour that cart abandoners on mobile between 8pm and 10pm respond better to free shipping than to 10% off, acts on it, and keeps re-checking as the pattern decays.

That difference changes how you read a vendor claim. Ask whether there is a model training on your data and updating over time, or a rules engine with AI vocabulary bolted on. Both can be worth buying. Only one of them degrades if you stop feeding it, and only one can surprise you.


Four core concepts every CMO must own

  1. TRAINING DATA: THE ONLY THING YOUR MODEL KNOWS

Training data is a table. Each row is one example: a customer, a session, an email send. Most columns are features, the signals available at the moment of the decision. One column is the label, the outcome you want to predict, recorded after the fact. The model learns the association between features and label. Nothing outside that table exists for it.

Two consequences that marketers underrate. First, the label defines the product. Train on "opened the email" and you have built an opener predictor, which will happily rank a subscriber who opens everything and buys nothing at the top. Train on "purchased within 14 days" and you get a different model with different winners. Nobody in the vendor demo will ask you which label you want; decide it yourself before procurement starts.

Second, base rates. If 2% of your subscribers churn in a quarter, a model that predicts "nobody churns" is 98% accurate and useless. Accuracy on rare outcomes is close to meaningless. The question to ask is how much richer the top-scored decile is than a random decile. As a rough floor on volume, you want thousands of labelled positive examples, and ideally not all from the same three months, because a model trained only on Black Friday behaviour has learned Black Friday.

  1. PROPENSITY MODELS: ONE NUMBER PER PERSON, PER DECISION

A propensity model is supervised learning applied to a marketing decision. It reads a customer's features and returns a probability between 0 and 1: likelihood to churn, to buy, to upgrade, to open. Lifetime value models work the same way but return an amount rather than a probability.

The output is a score, not an instruction. A score does nothing until you attach a threshold and an action, and choosing that threshold is a commercial decision about how much you will spend to save a customer, not a technical one. A propensity model also does not tell you why. It reports that this customer resembles people who left; it cannot tell you that your delivery times slipped in the north of the country. That remains a job for analysts.

  1. GENERATIVE MODELS: THEY PRODUCE CONTENT, NOT PROBABILITIES

Large language models (OpenAI's GPT series, Anthropic's Claude and their peers) are trained to predict the next fragment of text across enormous corpora. Point them at marketing and they draft subject lines, product descriptions, ad variants and chatbot replies. Image models do the same for visuals. Older natural language processing tooling sits alongside them for sentiment analysis on reviews and social posts.

The structural difference matters more than the capability. A propensity model has a label, so you can check whether it was right. A generative model has no ground truth for "good ad copy", so there is no accuracy figure, only human review or a downstream test. It also knows nothing about your customers unless you put customer data into the prompt or the retrieval layer, which is where the persistent profile described in the CDP foundations lesson does the work. Buying a generative tool and expecting propensity is one of the most common category errors in MarTech procurement.

One practical note: the model names move fast, and vendors quietly swap the model underneath their product. Ask which model version powers a tool today and how you are notified when it changes, because a swap can shift the tone and length of generated copy overnight.

  1. WHERE PREDICTION GENUINELY BEATS RULES

Prediction wins when the decision repeats thousands of times a day, the outcome is observed quickly and often, and the useful pattern involves interactions between more signals than anyone would write by hand. Rules win more often than AI vendors admit: when volume is low, when the outcome takes 14 months to observe as in long-cycle B2B, when the constraint is legal or ethical (never email someone who unsubscribed), and when someone must be able to explain a specific decision to a regulator or a customer.

Attribution is the clearest case of a rule being beaten. Last-click is a rule, and a bad one, because it hands all credit to the final touch. Data-driven attribution models the actual sequence of touchpoints and assigns fractional credit statistically. Google went past offering this as an option: it retired the rule-based models (last-click, first-click, linear, time-decay, position-based) in Google Ads and made data-driven attribution the standard. Note the condition, though. Statistical attribution needs conversion volume to separate signal from noise, which is exactly why it beats a rule for a high-volume advertiser and not for one with 40 conversions a month.


Real-world results from named brands

NETFLIX: In their 2015 paper on the recommender system, Netflix researchers put roughly 80% of streamed hours as coming from recommendations rather than search, and estimated the combined annual value of personalisation and recommendation at more than a billion dollars, mostly through retention. The instructive part is the Netflix Prize. Netflix paid $1m in 2009 for an algorithm that improved rating prediction by just over 10%, then said publicly that it did not put the full winning solution into production: the engineering cost of running that ensemble outweighed the accuracy gain. Model accuracy and business value are different currencies, and the exchange rate is often poor.

A second-order effect worth understanding: every row of a recommender's training data records what the system itself chose to show. Titles that were never surfaced generate no engagement, which looks to the next training run like a lack of interest. Left alone, the model narrows towards what it already promotes, which is why serious recommendation teams deliberately inject exploration.

SPOTIFY: Discover Weekly, launched in 2015, reached tens of millions of listeners within its first year. It blends collaborative filtering (people whose listening overlaps with yours also played this) with signals from the Echo Nest, the music intelligence company Spotify bought in 2014, which analysed audio characteristics and text written about artists across the web.

That blend exists because of the cold start problem, the clearest limit of learning from history. A track uploaded this morning has no listening history, so collaborative filtering has nothing to work with. Only a model that listens to the audio itself, or reads about it, can place a brand new release. The same gap opens for your first-time visitors: whatever propensity model you build, roughly the entire new-customer population arrives with no features worth scoring, and that segment needs rules or contextual signals instead.

How Netflix Uses AI and Machine Learning

Watch on YouTube

Knowledge check

1. What is the key distinction between Machine Learning and a traditional rules-based system in marketing?

2. Why does the lesson stress that a CMO must be able to tell whether a vendor's 'AI' is a rules engine or a real model?

3. In the lesson, what best describes 'supervised learning'?

MULTIPLE CHOICE

4. Select ALL statements that correctly reflect the lesson's view of AI and ML in marketing.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL tasks that are examples of supervised learning applications described or implied in the lesson.

Select all the correct answers.

🎬 note: the quiz above sits roughly two thirds through the lesson.


CMO action items

  • Require every vendor using the word AI to state in writing whether the system runs static rules or a model trained on your data, what label that model predicts, and which underlying model version powers any generative feature. That single set of questions removes half the noise immediately.
  • Write down the label before you commission anything. "Likely to buy" is not a label. "Placed an order within 30 days of the send" is. Argue about it with finance now rather than after the first campaign.
  • Commission a predictive lifetime value model for your largest segment within 90 days. You do not need an in-house ML team; Databricks, BigQuery ML and the built-in scoring in most CRM suites will run a first model on existing data, though note that all of these vendors sell the tooling they benchmark. Use the output to set different acquisition cost ceilings by segment.
  • Give one named person the feedback loop between model outputs and the campaign teams. Models drift when the market moves, and someone has to notice.

Common mistakes that kill results

Mistake 1: deploying ML on dirty data

A model is only as good as its table. Duplicate customer records split one person's history into three weak ones, and a model trained on those learns that everybody is a light buyer. Missing behavioural signals are worse than absent, because the model treats "not recorded" as "did not happen". Run identity resolution and a data quality audit before any modelling work, not after the first disappointing result.

Mistake 2: label leakage, the failure that looks like success

Leakage happens when a feature encodes the outcome. A churn model fed "number of contacts with the cancellation team" will look astonishing offline and be worthless in production, because by the time that feature exists the customer has already gone. Same for a purchase model that includes "discount code redeemed". Treat a suspiciously excellent test result as a smell rather than a triumph, and demand a list of features with the exact timestamp at which each becomes available.

Mistake 3: treating AI as a set-and-forget system

Behaviour shifts and the training table goes stale. Spotify described how listening in spring 2020 moved off commutes and onto home speakers, TVs and consoles, with weekday patterns resembling weekends. Any model using time of day, device or context as features inherited that shift the day it happened. Build monitoring and a retraining cadence into the operating plan from the start, and re-test your prompts and outputs whenever a generative vendor changes the model underneath. How to prove the resulting lift is real belongs to the methodology lesson; noticing that it has gone is your job every month.


Key takeaways

  • AI is the category; ML is the part that learns from your data and updates itself. Make vendors say which one they sell and which model version runs underneath.
  • Training data is a table of features plus one label, and the label you choose is the product you get. "Opened" and "bought" build different models with different winners.
  • On rare outcomes, accuracy is a vanity metric. A 2% churn base rate makes "nobody churns" 98% accurate. Ask how the top decile compares with random.
  • Propensity models return a probability per person and no explanation. Generative models return content and have no ground truth to be scored against. Do not buy one expecting the other.
  • Prediction beats rules at high volume with fast, frequently observed outcomes. Rules stay better for low volume, long sales cycles, and any decision you must explain to a regulator. Google Ads retiring rule-based attribution is the canonical case of the former.
  • Netflix paid $1m for a 10% accuracy gain it chose not to fully deploy. Accuracy and value are not the same currency.
  • Cold start is a permanent limit: new items and new customers arrive with no history, so plan rules or contextual signals for that population.
  • Dirty data, label leakage and set-and-forget deployment are the three failure modes that quietly waste the budget.

Resources

  • 🔗
    Google's Data-Driven Attribution Overview

    Google's official documentation explaining how data-driven attribution works in Google Ads and how it differs from last-click models, with guidance on switching and interpreting results.

  • 🔗
    Salesforce State of Marketing Report

    Annual benchmark report from Salesforce covering how marketing teams are actually deploying AI and ML tools, with adoption rates, use cases, and performance data from thousands of surveyed marketers.

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Run a data quality audit before any ML initiative
  • Define one 90-day ML business outcome and test against a holdout
  • Set model monitoring with retraining triggered when predictions drift beyond 10%
See the full action playbook →