Menu

Course · Data platforms

ETL and ELT

ETL transforms data before loading it; ELT loads first and transforms inside the warehouse or lakehouse. Learn when each fits.

Lessons
8
Interview questions
1
Projects & case studies
2
Reading time
~1 h

About this course

ETL and ELT describe where transformation happens. In ETL it happens in a separate processing step before data reaches the target. In ELT raw data is loaded first and transformed inside the warehouse or lakehouse with SQL.

Neither is universally better. The choice depends on data volume, where compute is cheapest, governance needs and team skills.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. ETL vs ELT and Modern Data PipelinesHow modern data pipelines are built: ETL vs ELT, batch vs streaming, layered models, orchestration, idempotency, data quality, observability and recovery.Beginner2 min

Beginner

Core concepts you will use every day.

  1. ETL vs ELT: Choosing the Right ApproachETL transforms data before loading it, ELT loads raw data first and transforms inside the warehouse. Learn the factors that decide which one fits your pipeline.Beginner3 min

Intermediate

Patterns used in production pipelines.

  1. Batch vs Streaming Data PipelinesWhen to process data in batches and when to stream it: latency needs, complexity, correctness with late data, cost, and the micro-batch middle ground.Intermediate2 min
  2. Idempotency in Data PipelinesMake every pipeline step safe to rerun: deterministic inputs, overwrite and merge patterns, atomic publishing, idempotent consumers and guarded side effects.Intermediate2 min
  3. Data Quality: Checks, Contracts and Failure HandlingBuild a data quality practice: which checks to run where, how contracts set expectations between teams, and how to decide whether a failure blocks or alerts.Intermediate2 min
  4. Data Pipeline Observability FundamentalsWhat to measure and alert on in data pipelines: run status, duration, volume, freshness, quality results, lag and lineage, so problems are found before consumers notice.Intermediate2 min
  5. Data Pipeline Reliability and Retry DesignDesign pipelines that recover on their own: classify failures, retry with backoff and limits, isolate bad records, make steps idempotent and plan backfills.Intermediate2 min

Advanced

Performance, internals and edge cases.

  1. CDC Patterns and Failure ModesCompare change data capture approaches (log-based, query-based, triggers, outbox) and the failure modes to design for: ordering, deletes, snapshots and schema changes.Advanced2 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type