ETL and ELT courseLesson 6 of 8
ETL and ELT course · Lesson 6 of 8
Data Pipeline Observability Fundamentals
What to measure and alert on in data pipelines: run status, duration, volume, freshness, quality results, lag and lineage, so problems are found before consumers notice.
On this page
A pipeline can succeed and still produce wrong or late data. Observability means you can tell, from signals you already collect, whether data is arriving on time, in the right amount and with the right content, and where to look when it is not.
The signals
| Signal | Question | Example metric |
|---|---|---|
| Run status | Did it run and succeed? | Success/failure per task |
| Duration | Is it getting slower? | Runtime per task vs recent baseline |
| Volume | Is the amount plausible? | Rows in, rows out, rows rejected |
| Freshness | Is the data current? | Time since last successful publish per table |
| Quality | Is it correct? | Check results per run |
| Lag (streaming) | Are we keeping up? | Consumer lag, end-to-end latency |
| Lineage | What is affected? | Upstream and downstream dependencies |
Alert on outcomes, not just failures
The most important alert is freshness: “table X has not been updated in N hours”. It catches failures, stuck runs, skipped schedules and silent upstream outages, all with one rule. Add volume and quality alerts next. Task-failure alerts are necessary but not sufficient.
Make runs self-describing
Write a small audit record for every run: pipeline, data interval, start and end time, rows read, written and rejected, check results and code version. That one table answers most incident questions and feeds dashboards and baselines.
Structured logs
Log with context (run id, data interval, table, file, row counts) in a consistent, parseable format, so you can search for “everything that happened to 2026-10-01”.
Lineage
When a table is wrong, lineage tells you which upstream table or job to inspect and which downstream dashboards to warn. Many catalogs and orchestrators capture lineage automatically; make sure yours does for production pipelines.
Common mistakes
- Monitoring only task failures.
- No record of row counts per run.
- Alerts with no owner or runbook.
- Too many alerts, all ignored.
Key takeaway
Measure status, duration, volume, freshness, quality and lag per run, alert first on freshness, and keep an audit record that explains every run.
Progress is saved in this browser only. No account needed.