Glossary
AIDatageneral

MLOps

Also: MLOps, Machine Learning Operations, ML Ops, LLMOps, ModelOps

Machine Learning Operations: combining ML and DevOps practices to industrialise, deploy, monitor, and retrain models reliably in production.

What it is

MLOps (Machine Learning Operations) is the set of practices, tools, and organisational disciplines used to take machine learning models from experimentation to reliable, repeatable production. It applies the automation and rigour of DevOps to the specific challenges of ML: data pipelines, model training, versioning, deployment, and continuous monitoring.

Unlike traditional software, an ML system depends on data as well as code. A model that worked last quarter can silently degrade as customer behaviour, prices, or market conditions shift. MLOps exists to manage this moving target.

Why it matters

Most ML projects never reach production, and many that do fail quietly. MLOps addresses the gap between a promising prototype and a dependable business capability.

  • Reliability: models behave predictably and can be rolled back.
  • Reproducibility: any result can be traced to specific data, code, and parameters.
  • Speed: new versions ship in hours or days, not months.
  • Governance: audit trails, access control, and documented lineage support compliance.
  • Cost control: automated retraining and monitoring prevent expensive silent failures.

How it is used in practice

A mature MLOps setup typically covers:

  • Data and feature pipelines: versioned, validated, and repeatable.
  • Experiment tracking: recording parameters, metrics, and artefacts.
  • Model registry: a catalogue of approved model versions and their status.
  • CI/CD for models: automated testing, packaging, and deployment.
  • Monitoring: tracking latency, accuracy, and data drift (inputs changing over time).
  • Automated retraining: triggered by schedule or by detected drift.

These practices apply equally to classic predictive models and to applied LLM systems (where prompts, retrieval sources, and evaluation sets also need versioning and monitoring, sometimes called LLMOps).

A concrete worked example

A bank deploys a credit scoring model.

1. Data scientists train a model; the run is logged with its dataset version and metrics.

2. The model passes automated fairness and accuracy tests, then enters the registry as "staging".

3. After human approval it is promoted to "production" and served via an API.

4. Monitoring flags that applicant income distributions have shifted (drift), and approval rates are climbing.

5. An automated pipeline retrains on fresh data, the new version is tested, and it replaces the old one with full audit history.

Without MLOps, step 4 might go unnoticed until losses or a regulator surface the problem.

The MLOps lifecycle Data and features Train and track Registry and deploy Serve in production Monitor and detect drift Retrain (automated)
MLOps as a closed loop: data and training feed deployment, while monitoring and drift detection trigger automated retraining.

Frequently asked questions

What is MLOps in simple terms?

MLOps (Machine Learning Operations) is the set of practices, tools, and organisational disciplines that take machine learning models from experimentation to reliable production. It applies DevOps automation and rigour to problems specific to ML: data pipelines, model training, versioning, deployment, and continuous monitoring. The core difference with classic software is that an ML system depends on data as well as code, so a model that worked last quarter can degrade silently as behaviour or market conditions shift.

What is the difference between MLOps and DevOps?

DevOps industrialises code; MLOps industrialises code plus data plus models. A traditional application behaves the same until someone changes the code, whereas an ML model can lose accuracy without any code change because the incoming data has shifted. That is why MLOps adds data and feature pipelines, experiment tracking, a model registry, drift monitoring, and automated retraining on top of standard CI/CD.

What is data drift and why does it force retraining?

Data drift means the inputs a model receives in production change over time compared with the data it was trained on. Since the model learned relationships from the old distribution, its predictions become progressively less reliable, often without any visible error. Monitoring for drift is what triggers retraining, either on a schedule or automatically when a shift is detected.

What does a mature MLOps setup actually include?

Six building blocks: versioned and validated data and feature pipelines, experiment tracking (parameters, metrics, artefacts), a model registry listing approved versions and their status, CI/CD for automated testing and deployment of models, monitoring of latency, accuracy and data drift, and automated retraining triggered by schedule or drift. Together they deliver reliability, reproducibility, faster releases, audit trails for governance, and protection against costly silent failures.

Do MLOps practices apply to LLM projects, or is that something else?

They apply, and the variant is sometimes called LLMOps. The same discipline holds for applied LLM systems, except that prompts, retrieval sources, and evaluation sets also need versioning and monitoring alongside the model itself. The underlying logic is unchanged: trace every result to specific inputs, test before promoting, and monitor after deployment.