Delta Lake courseLesson 1 of 6
Delta Lake course · Lesson 1 of 6
Data Lakes, Lakehouse Architecture and Delta Lake
From data lakes to the lakehouse: file formats, table formats, Delta Lake transactions, schema evolution, medallion layers, maintenance and governance in one guide.
On this page
A data lake stores files cheaply in object storage. A lakehouse adds the reliability of a warehouse on top of those files through an open table format such as Delta Lake. This guide connects the pieces.
1. Lakes, warehouses and lakehouses
Lakes are flexible and cheap but lack transactions and enforced schemas; warehouses are reliable but keep data in their own storage; lakehouses keep data in open formats on object storage while adding transactions and governance.
Read: Lake vs warehouse vs lakehouse · Practise: When to choose each
2. File formats
Use columnar Parquet for analytical data, row-based Avro for messages, and treat JSON and CSV as landing formats only.
Read: Parquet vs Avro vs ORC · JSON vs Parquet
3. Table formats and the transaction log
Delta Lake writes Parquet data files plus a _delta_log of commits. That log provides atomic writes, consistent snapshots, optimistic concurrency, time travel and row-level MERGE, UPDATE and DELETE.
Read: Delta Lake transactions, schema evolution and time travel · Delta vs plain Parquet tables · Practise: What Delta Lake solves
4. Schema enforcement and evolution
Reject mismatched writes by default; evolve schemas deliberately, preferring additive changes and expand-and-contract migrations for breaking ones.
Read: Schema evolution patterns · Practise: When schema evolution is safe
5. Medallion layers
Organise tables into bronze (raw), silver (cleaned and conformed) and gold (business models). Each layer reads only the previous one and can be rebuilt.
Read: Databricks workspace, jobs and lakehouse
6. Maintenance
Compact small files, cluster by common filters, and VACUUM old files with a retention period that respects time travel and long readers.
Read: Partitioning, clustering and data layout
7. Governance
A catalog provides names, permissions, masking and lineage across engines.
Read: Unity Catalog and governance
Design it
Put it together in the scalable lakehouse case study, then build the Kafka → Spark → Delta streaming project.
Progress is saved in this browser only. No account needed.