DataAI & ML Strategy

The ML-to-production field guide: who shapes what breaks and why

Most ML projects die somewhere between a promising notebook and a live system. This field guide identifies the people, projects, and organisations whose work reveals exactly where the gaps are and what serious practitioners do about them.

Listen to the podcast

4 min

Chapters

Key takeaways

  • Before approving any model, ask who by name is on the hook to keep it running six months after launch.
  • Budget for an ML engineer who sits between the data scientist and the software team, not just modellers.
  • Treat each dataset as a product with an owner, a quality standard, documentation and a promise it will not silently change.
  • Monitor for data drift, because a degrading model keeps returning confident answers instead of crashing.
  • Discount vendor productivity claims such as DBT Labs surveys and cross check them against O'Reilly Radar or MIT Sloan research.
Read the full transcript

Host:Welcome to the Leaders Insights Podcast. Today's episode, the ML to Production Field Guide, who shapes what breaks and why.

Expert:Here's a stat that should terrify anyone funding an AI budget. Most machine learning projects never make it out of the lab. They die somewhere between a data scientist notebook and an actual working product. Why? Because a notebook is a sandbox and production is a war zone. In the notebook, that's the interactive coding environment where data scientists prototype. Everything is clean, static, and forgiving. In production, the data is late. The users are weird. And nobody warned you the upstream system changed a column name at 2 a.m. So, it's not that the models are bad, it's everything around them. The model is maybe 10% of the work. The other 90 is plumbing. Pipelines that move data, monitoring that tells you when it rots, and the boring governance of who owns what. O'Reilly's survey work has been saying this for years. The top barriers to ML adoption aren't algorithms, they're skills gaps and lack of data infrastructure.

Host:Give me the concrete failure. Where does the notebook to production handoff actually snap?

Expert:Three places. One, the data drifts. The real world inputs slowly stop looking like the training data, so a fraud model trained on 20-24 spending patterns quietly goes senile. Two, nobody owns the pipeline after launch, so it breaks and sits broken. Three, the thing was never designed to be retrained, so fixing it means rebuilding it.

Host:Go senile. That's the drift problem. How fast does that happen?

Expert:Faster than people expect. A recommendation model can degrade meaningfully in weeks if behavior shifts. And here's the trap. It doesn't crash. It keeps returning confident answers that are increasingly wrong. A broken server pages you at midnight. A drifting model just costs you money silently for a quarter.

Host:So who are the people who actually fix this?

Expert:The field guide is about roles. The key hire nobody budgets for is the ML engineer. The person who lives between the data scientist and the software team, whose whole job is making a model survive contact with reality. Then data engineers who build the pipes. Use those two and you've hired a Formula One driver and forgotten to build a track.

Host:Let's talk numbers. I've seen vendors claim huge productivity gains from better tooling. How much do you trust that?

Expert:Carefully. DPD labs. And note. They sell data transformation software, so they've got skin in the game. Publishes analytics engineering surveys showing data teams spend enormous time just cleaning and reshaping data before any modeling happens. Directionally true, but cross check it against something independent. MIT Sloan Management Review has done more sober work here and their line is consistent. The winners treat data as a managed product, not a byproduct.

Host:Data as a product. Unpack that because it sounds like a slogan.

Expert:It means someone owns a dataset the way a product manager owns an app. With a quality standard, documentation, and a promise it won't silently change. Most companies treat data like leftovers in the office fridge. Technically available, ownership unclear, increasingly dangerous.

Host:So an exec hears all this and thinks, I need to buy a platform. Are they right?

Expert:That's the reflex, and it's usually wrong. Tools don't fix an org that hasn't decided who's accountable when the model breaks at 3 a.m. MIT Sloan's research keeps landing on culture and clear ownership over any specific tool. You can buy the best monitoring in the world and still ignore the alerts.

Host:Give me the uncomfortable question then. What's the one thing companies fake?

Expert:They fake done. A demo works, everyone claps, the press release goes out. And then it never actually runs reliably for real users. Working in the demo and working in production are as different as a wedding and a marriage.

Host:So for the exec listening who's about to greenlight an AI project, what do they do Monday morning?

Expert:Before you approve the model, ask one question. Who is on the hook to keep this running six months after launch? And what's their name? If there's no name, you're not funding a product. You're funding a very expensive demo.

Host:A name, not a platform. That's the takeaway. Thanks for this.

Expert:Anytime.

Host:Sources for today's episode. Towards Data Science, KD Nuggets, The New Stack, DBT Labs, Vendor, Data Tooling, O'Reilly Radar, MIT Sloan Management Review. That's a wrap. Watch CDO briefings drop daily at MBA-training.com.

The shortlist below is grouped by influence on the practitioner community, meaning the degree to which each entry has changed how data teams actually think about moving models from experiment to production. Revenue figures are irrelevant here; what matters is whether teams at real companies have changed their behaviour because of this work. Five entries, no filler.

Five entries reshaping the path to production ML

DuckDB and DuckLake, a lakehouse without a cluster

DuckDB started as an academic project at CWI Amsterdam and has become the runtime of choice for engineers who need analytical SQL without spinning up a cluster. The recent DuckLake work, covered in Towards Data Science, extends the idea: start with a local Parquet file, join it to data stored in the cloud, and you have a functional lakehouse in minutes. What makes this notable for production ML is the friction it removes at the data layer. Most pilots fail not because the model is wrong but because the data pipeline is a mess of undocumented dependencies. DuckLake gives small teams a path to aunified analytical and ML data layer without a six-month infrastructure project first. That changes the pilot-to-production calculus materially.

dbt State and the open lakehouse pivot at dbt Labs

At dbt Summit 2026, dbt Labs (a commercial data tooling vendor, so treat the framing with appropriate scepticism) announced dbt v2, dbt State, and a stated "open lakehouse vision" alongside Fivetran. dbt State is now generally available and the company's own description is that it was built to stop rebuilding what had not changed: the system tracks transformation state so only modified models rerun. For production ML, this matters because retraining pipelines routinely rebuild everything regardless of what changed, burning compute and introducing unnecessary variance. The dbt State approach, if applied to feature pipelines, is a practical answer to that waste. The commercial relationship between dbt Labs and Fivetran means the full vision should be evaluated against independent benchmarks before committing, but the underlying mechanism is sound.

Zed's Delta and agent-native deployment workflows

Zed's Delta project, reported by The New Stack in September 2026, launched on the premise that AI agents have made the pull request model obsolete and that everyone is racing to replace GitHub's collaboration layer. This is relevant to ML production for a less obvious reason: the bottleneck in most MLOps pipelines is not model quality but code review and deployment ceremony. When agents generate, test, and merge code faster than human review cycles allow, the organisational process becomes the constraint. Zed's bet is that the tooling layer needs to rebuild around agent-native workflows. Whether Delta succeeds commercially is a separate question; the diagnosis of where production ML slows down is accurate.

O'Reilly Radar names the query gap in enterprise analytics

O'Reilly's ongoing Radar work on intelligent data orchestration with LLMs documents a pattern that anyone who has tried to ship a model to production will recognise: organisations spend years building dashboards that answer last quarter's questions, then discover that the model they trained needs answers to this week's questions. The observation from O'Reilly's analysts, drawn from 17 years of enterprise platform work, is that the question "can I ask one thing and get one answer across everything my company knows" has driven every serious data architecture decision. The production ML failure mode this names is the query gap: the training distribution does not match the operational distribution, anddetecting that drift before it reaches your users requires infrastructure most teams have not built.

MIT Sloan on the skills lag between pilot and production

MIT Sloan's 2026 research on emerging tech skills found that reskilling programs are typically built around forecasted skill needs rather than observed ones, which creates a predictable mismatch. For CDOs trying to move pilots into production, this surfaces as a hiring and team composition problem: the person who built the pilot in a notebook often does not have the software engineering background to own the production system, and the software engineer assigned to productionise the model often does not understand what the model is doing. MIT Sloan's data suggests the gap is structural, not a training problem that a weekend course fixes. That is an uncomfortable finding, and it is worth taking seriously precisely because it comes from independent research rather than a vendor with a platform to sell.

What do these five standouts have in common?

Every entry on this list is responding to the same underlying problem from a different angle. DuckDB and DuckLake attack the data infrastructure gap. dbt State attacks the compute and reproducibility gap. Zed attacks the deployment workflow gap. O'Reilly names the query distribution gap. MIT Sloan names the skills gap. None of them is solving "ML." They are each solving one specific failure mode in the path from a working model to a working system.

The implication for CDOs is that "ML to production" is not a single problem with a single solution. It is five or six distinct failure modes that happen to be sequential. Organisations that treat it as one problem tend to fix the most visible one, usually model accuracy, and then discover the next one waiting. The teams making genuine progress in 2026 are the ones that have named each failure mode separately and assigned ownership accordingly.

The honest lesson from this shortlist: the most valuable skill in ML production is not building better models, it is knowing which layer is actually broken.

Who to watch: the small engineering teams combining DuckDB-native data layers with agent-assisted deployment pipelines, because they are collapsing three of these failure modes simultaneously and doing it without enterprise budgets.

Go deeper

The lessons that take this article further, free to read.

  1. 1Closing the POC-to-production gapAI & machine learning strategy
  2. 2Models in production: drift, monitoring & MLOpsAnalytics, BI & decision intelligence
  3. 3MLOps: monitoring, retraining & driftAnalytics, BI & decision intelligence
  4. 4The lakehouse: unifying analytics & MLModern data architecture
  5. 5Data observability: detect problems before your usersModern data architecture

Finished reading?

Validate your read to earn XP and feed your radar.