Roadmap
Data Engineer Roadmap: Beginner to Job Ready
An ordered Data Engineering roadmap from foundations to job ready: SQL, Python, warehousing, ETL/ELT, PySpark, cloud, streaming, system design, projects and interviews.
How to use this roadmap
This roadmap is a learning map, not a checklist of everything. Work through the stages in order; each one builds on the last. Time estimates assume roughly 8 to 12 focused hours a week and vary a lot with prior experience, so treat them as guidance, not targets.
Skip a stage only if you can already meet its outcome. Spend your time on the outcomes, not on finishing pages.
The journey, stage by stage
Stage 1: Beginner foundations
Command line, Git, how data moves between systems, and basic data formats such as CSV, JSON and Parquet.
- Why it matters
- Every later tool assumes you are comfortable with files, terminals and version control.
- Typical effort
- 2–3 weeks
- Outcome
- You can navigate a terminal, version your work with Git and explain the main file formats.
Stage 2: SQL
Joins, aggregation, window functions, CTEs, and reading a query plan.
- Why it matters
- SQL is used in every warehouse, lakehouse and Spark SQL job.
- Typical effort
- 4–6 weeks
- Outcome
- You can write correct multi-table queries and explain row counts after a join.
Stage 3: Python
Functions, modules, error handling, logging, generators and testing data code.
- Prerequisites
- SQL basics
- Typical effort
- 4–6 weeks
- Outcome
- You can build a small loader that is safe to run twice and has tests.
Stage 4: Data warehousing
Dimensional modelling: facts, dimensions, grain, star schemas and slowly changing dimensions.
- Prerequisites
- SQL
- Typical effort
- 2–3 weeks
- Outcome
- You can design a star schema for a business process and state its grain.
Stage 5: ETL / ELT
Where transformation runs, layered models (raw, staging, marts), idempotent loads and data-quality checks.
- Prerequisites
- Data warehousing
- Typical effort
- 2 weeks
- Outcome
- You can explain ETL versus ELT and build a rerunnable load.
Stage 6: PySpark
DataFrames, joins, window functions, partitions, shuffles and skew.
- Prerequisites
- SQL, Python
- Typical effort
- 4–6 weeks
- Outcome
- You can write a PySpark job and explain where its shuffles happen.
Stage 7: Cloud
Object storage, compute, identity and access, and cost basics on one major cloud (AWS, Azure or GCP).
- Prerequisites
- Python
- Typical effort
- 3–4 weeks
- Outcome
- You can store data in object storage and run a job against it with least-privilege access.
Stage 8: Databricks / Snowflake
A lakehouse or cloud warehouse platform, and table formats such as Delta Lake.
- Prerequisites
- PySpark or SQL, Cloud
- Typical effort
- 3–4 weeks
- Outcome
- You can explain how your chosen platform stores data and scales compute.
Stage 9: Kafka / Airflow
Event streaming with Kafka and pipeline orchestration with Airflow.
- Prerequisites
- Python
- Typical effort
- 3–4 weeks
- Outcome
- You can explain partitions and consumer groups, and schedule an idempotent DAG with retries.
Stage 10: Data Engineering system design
Requirements, scale, architecture, reliability, data quality and trade-offs for data platforms.
- Prerequisites
- Most earlier stages
- Typical effort
- 3–4 weeks
- Outcome
- You can walk through a batch or streaming design and defend its trade-offs.
Stage 11: Projects
Two or three end-to-end projects you can run, test and explain.
- Typical effort
- 4–8 weeks
- Outcome
- You have projects with tests, data-quality checks and a clear story.
Stage 12: Interview preparation
Technology questions, system design practice, behavioural stories and company preparation.
- Typical effort
- 3–6 weeks
- Outcome
- You can answer core questions concisely and handle follow-ups.
Stage 13: Job ready
Targeted applications, a resume built on real results, and a steady interview routine.
- Typical effort
- Ongoing
- Outcome
- You apply with evidence: projects, explanations and honest results.
This site does not save personal progress. Use the stages as an ordered guide and keep your own checklist.