+100 XP

AI & ML in marketing: real-world application

In 2018 L'Oréal bought ModiFace, the Toronto company whose augmented-reality try-on worked from facial-geometry mapping rather than static shade charts. That purchase was a decision to own the model layer instead of renting it from an agency or a platform. Follow the programme from those first diagnostics through to the assistant L'Oréal showed at CES 2024 and you get an unusually complete picture of what AI personalisation costs a global beauty group, where it snaps, and which numbers survive contact with a CFO.

Core concept: what l'oréal actually put into production

Three pieces, built separately and later stitched together. A diagnostic that reads a selfie and returns skin or hair signals. A try-on renderer that places a shade on a live face. A recommender that maps those outputs plus purchase history onto specific SKUs. All of it sits on consented profiles of the kind the CDP foundations lesson describes, and each brand runs its own version tuned to its catalogue, because Vichy's dermatological range and Maybelline's mass colour range are not the same recommendation problem. At group scale (36 brands, roughly 150 countries, €41.2 billion of sales in 2023) that means dozens of things in production, not one.

Key sub-concept 1: the pilot was a diagnostic, not a campaign

L'Oréal did not start with a personalised email programme. It started with Vichy SkinConsult AI in 2019, a selfie diagnostic developed with dermatologists that grades signs of ageing and returns a regimen, followed by La Roche-Posay's Effaclar Spotscan for acne severity. That choice is the interesting part. Dermatologist-graded image sets give you labelled ground truth, which almost no marketing dataset has. The objective was narrow, the output was checkable against a clinician, and there was exactly one funnel step to watch: diagnostic result to basket. When the pilot underperformed, you knew whether it was the grading or the product mapping. A pilot on open-ended "personalised content" gives you neither.

Key sub-concept 2: what broke between pilot and rollout

Four things, in roughly this order.

Coverage of the training set. Beauty AI trained mostly on well-lit images of lighter skin tones degrades at the extremes of the shade range, and the extremes are exactly where mismatched foundation gets returned. Fixing this is a data acquisition project with a cost line, not a model tuning exercise.

Camera variance. A phone's automatic white balance rewrites the colour of the input. A diagnostic that is accurate in a studio and unstable under kitchen lighting produces different recommendations for the same face on two consecutive evenings, and consumers notice.

Consent. Illinois BIPA treats facial geometry as biometric data requiring informed written consent before capture, and beauty try-on tools have attracted class actions on exactly that basis. The result is a jurisdiction-aware gate in front of the experience, plus a structural hole in the dataset: the users who decline are not random, so the model learns from a skewed population.

Distribution you do not own. ModiFace-powered try-on ran on partner surfaces including Amazon and Google, which is how the technology reached consumers fastest. It also means the interaction signal stays with the retailer. You buy conversion there; you buy learning only on your own properties. The second-order consequence is that the model improves fastest on the channels with the least traffic, and a rollout plan has to budget for that asymmetry rather than discover it in year two.

Key sub-concept 3: what was measured, and the comparison that lies

The tempting number is conversion rate among try-on users versus everyone else. It is always spectacular and it is close to meaningless: people who open a try-on module are already deciding. Any honest read has to randomise access to the module or hold out geographies, using the discipline the frameworks lesson sets out for uplift.

The line that matters in beauty is returns. Cosmetics returns are frequently unsellable once opened, so a returned foundation carries the reverse logistics cost plus a full write-off of the unit. On a category where online return rates sit in the double digits, taking two points off returns through better shade matching is usually worth more margin than two points of conversion, and it is easier to attribute because the SKU and the reason code are both recorded. Alongside it, track consent opt-in rate by market, accuracy by skin tone decile rather than in aggregate, and repeat purchase interval for diagnosed customers.

How Google's AI Powers Smart Bidding

Watch on YouTube

Key sub-concept 4: the generative layer landed on top of all this

At CES 2024 L'Oréal presented Beauty Genius, a generative assistant that answers beauty questions in free text and routes to products. This changes the risk profile rather than the plumbing. A recommender can only pick from a catalogue; an assistant writes sentences, and a sentence about acne, hair loss or eczema can cross into regulated claim territory in the EU and the US. The governance question belongs to the CMO playbook lesson. The commercial upside is quieter: free-text queries are a better statement of intent than any clickstream, which makes the assistant log the most valuable first-party dataset in the group, and also the one with the shortest legal fuse if retention and consent are sloppy.

Real-world case: the rollout arithmetic

What turned a set of brand experiments into infrastructure was the 2020 demand shock: L'Oréal's e-commerce grew more than 60% that year and reached roughly a quarter of group sales. Volume is what makes personalisation pay. Run the arithmetic on your own numbers before you fund anything. A diagnostic that lifts basket conversion by one point on 50,000 monthly sessions and a €60 average order returns about €30,000 a month, which will not cover a data science team, image licensing and a legal review in nine jurisdictions. The same one point on 20 million sessions is a different business. L'Oréal could justify owning the model layer because it had the traffic and the SKU count to amortise it across 36 brands. A single-brand retailer with a tenth of the traffic is usually better served buying the capability and spending its own budget on the input data.

Real-world case: the part that stayed niche

The programme also produced hardware. Perso, shown at CES 2020, was an at-home device that mixed personalised skincare and lipstick formulas, and it reached market as YSL Rouge Sur Mesure at a premium price point. Compare its reach with the try-on module, which appeared anywhere L'Oréal or a retail partner had a page. Software personalisation scales at close to zero marginal cost; the same intelligence expressed as a device inherits a bill of materials, cartridge logistics, firmware support and a returns policy for electronics. Both were real deployments of the same models. Only one of them scaled, and the difference had nothing to do with model quality.

Stitch Fix's Data Science Approach to Fashion

Watch on YouTube

CMO action items

  • Pick a first use case where an outside expert can grade the output, the way dermatologist scoring anchored the Vichy diagnostic. If nobody can say whether a prediction was right, you will spend a year arguing about dashboards.
  • Before any pilot, write down what you will do when the model is measurably worse for one customer group than another. Coverage gaps in beauty show up as wrong shades; in credit or insurance they show up as something a regulator reads.
  • Map which of your AI-powered surfaces you control and which sit on a retailer or platform. Negotiate signal return in the partnership contract, not after launch.
  • Run the volume arithmetic against a build-versus-buy comparison. Owning models makes sense at L'Oréal's traffic and catalogue depth; below some threshold it is a vanity project with a payroll.

Common mistakes that kill results

  • Reporting engaged-user versus everyone-else comparisons to the board. Self-selection will hand you a 3x conversion figure you cannot reproduce under a holdout, and you only find out when someone asks the model to forecast next quarter.
  • Optimising for the wrong ledger line. A recommender pushed towards conversion will happily sell the wrong shade twice, generating a return, a write-off and a customer who stops trusting the diagnostic. Set the objective against contribution margin after returns.
  • Treating consent as a legal checkbox rather than a data input. A declined-consent population is a hole in your training set, not just a missed session, and in biometric-adjacent categories it can be a large one.
  • Assuming that because the pilot worked in one market it will hold in the next. Camera hardware, lighting, skin tone distribution, claim rules and consent regimes all change at the border, and every one of them touches the model.

Resources

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Run a data quality audit before any ML initiative
  • Define one 90-day ML business outcome and test against a holdout
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.