+55 XP

Data observability: detect problems before your users

Data observability is the practice of understanding the health of your data, continuously, automatically, and proactively. It borrows from software engineering's observability (logs, metrics, traces) and applies it to data pipelines and datasets.

Without observability, you learn about data quality problems from your business stakeholders, which means the problem has already made it to production. With observability, you detect problems in the pipeline before they reach consumers.

The five pillars of data observability

1. Freshness, Is your data up to date? How recently was a table updated? Is the update cadence consistent with expectations? Freshness violations (a table that normally updates hourly but hasn't updated in 6 hours) are often the first indicator of a pipeline failure.

2. Volume, Is the expected amount of data arriving? A table that normally receives 10,000 rows per hour receiving only 100 rows is a signal, source system failure, upstream pipeline issue, or data loss.

3. Distribution, Are the statistical properties of your data consistent? If 5% of order amounts are normally negative (refunds), a sudden jump to 30% negative suggests data corruption. ML-powered anomaly detection identifies distribution shifts automatically.

4. Schema, Did the schema of your data change unexpectedly? A source system adding or removing a column without notice breaks downstream consumers. Automated schema change detection prevents silent failures.

5. Lineage, When a problem is detected, can you trace it upstream to the source and downstream to affected consumers? Lineage turns a "data is wrong" alert into "this table is wrong, it affects these 3 dashboards, and the root cause is this source system."

Data Observability Explained

Watch on YouTube

Knowledge check

1. What is the fundamental value proposition of data observability compared to a situation without it?

2. A table that normally updates every hour hasn't been updated in 6 hours. Which pillar of data observability does this most directly relate to?

3. Why is lineage considered especially valuable when a data problem is detected?

MULTIPLE CHOICE

4. Select ALL statements that correctly describe the pillars of data observability.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct statements about the data observability tooling landscape.

Select all the correct answers.

The data observability tooling landscape

Monte Carlo, The market leader. Full-suite observability: freshness, volume, distribution monitoring, lineage, automated anomaly detection. Enterprise pricing. Best for: large data platforms with complex pipelines.

Soda, SQL-based data quality checks integrated into pipelines. Define checks in YAML, run in dbt, Airflow, or any orchestration tool. More DIY than Monte Carlo but more flexible for custom use cases.

Great Expectations, The open-source standard for data validation. Define "expectations" (the data should have X% non-null for this column, values should be in this range). Generates data quality reports.

Elementary, Open-source, dbt-native. Generates observability metrics directly from dbt model runs. Free and easy to adopt for teams already on dbt.

Bigeye, Similar to Monte Carlo, focused on automatic anomaly detection without manual threshold configuration.

Choosing between them: Monte Carlo for enterprise scale and budget, Elementary + Great Expectations for cost-sensitive teams on dbt, Soda for teams wanting SQL-defined checks with code-level control.

Building observability into your pipeline

Observability shouldn't be bolted on after the fact, it should be designed into each stage of the pipeline.

At ingestion: Check that expected records arrived, schema matches the contract, and no critical fields are null.

At transformation: dbt tests run after each model materialization. Freshness checks alert if models don't run on schedule.

At serving: Monitor query patterns for unexpected null rates, row count anomalies, or metric deviations.

A mature implementation creates an "observability dashboard", a single pane of glass showing the health of all data products, updated in real-time, with drill-down to affected tables and upstream root causes.

The economics of data quality

McKinsey estimates that poor data quality costs the average Fortune 1000 company $15-25 million annually in operational inefficiencies, wrong decisions, and remediation work. Data observability has a measurable ROI.

Quantify it for your organization: count the analyst hours spent debugging data issues monthly. Multiply by hourly cost. Add the opportunity cost of wrong decisions made on bad data. That's your observability investment justification.

Quiz Questions

  1. Lequel de ces 5 piliers de l'observabilité des données détecte les changements statistiques inattendus dans les données ?

A) Freshness

B) Volume

C) Distribution

D) Schema

Réponse: C

  1. Quelle est la principale différence entre Monte Carlo et Great Expectations ?

A) Monte Carlo est gratuit, Great Expectations est payant

B) Monte Carlo est une plateforme enterprise full-suite avec ML, Great Expectations est open-source avec des validations définies manuellement

C) Great Expectations supporte plus de sources de données

D) Monte Carlo ne supporte pas le lignage de données

Réponse: B

  1. À quel moment l'observabilité des données doit-elle être intégrée dans un pipeline ?

A) Uniquement à la fin, avant la livraison aux consommateurs

B) Seulement lors des incidents de qualité

C) À chaque étape du pipeline, ingestion, transformation et serving, dès la conception

D) Uniquement dans les environnements de production

Réponse: C

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Embed data observability at every pipeline stage from design
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.