data-engineering interview questionsQuestion 2 of 6
Data Engineering interview question · Question 2 of 6
Data lake vs warehouse vs lakehouse: when would you choose each?
Short answer
I would use a data lake when I need cheap, flexible storage for raw or unstructured data that many engines can read, a warehouse when the main users are analysts and BI tools that need reliable, governed SQL with little operational effort, and a lakehouse when I want both: lake storage with transactions, schema enforcement and time travel through an open table format such as Delta Lake or Iceberg, shared by BI and ML workloads. The deciding factors are users, data types, openness, operations and cost.
Detailed explanation
Anchor the answer on properties:
- Transactions and schema enforcement: warehouse or lakehouse.
- Raw and unstructured data at low cost: lake (often as the bronze layer of a lakehouse).
- Multiple engines over one copy: lakehouse with an open table format.
- Minimal operations for SQL users: managed warehouse.
A typical combined design
Raw files land in object storage, curated tables are maintained in an open table format, and analysts query them through a SQL endpoint or warehouse. The table format is the source of truth; no data is copied without a reason.
Common mistakes
- Reciting definitions without tying them to the scenario’s users and data.
- Ignoring governance and ownership, which is what turns a lake into a swamp.
Progress is saved in this browser only. No account needed.