RoadmapsPlan 2 of 9
Study plan · Plan 2 of 9
30-Day Data Engineering Fundamentals Plan
A four-week plan that covers SQL, Python, data modelling and an idempotent pipeline, with a concrete outcome to reach at the end of each week.
This plan compresses the first stages of the Data Engineer roadmap into four weeks. Each week has one outcome. Do not move on until you can meet it; slipping a week is better than skipping understanding.
How to use it
- Block the hours in your calendar before the week starts.
- Write your answers and code in a Git repository so you can show your work later.
- At the end of each week, explain that week’s outcome out loud in two minutes. If you cannot, repeat the hardest part.
The plan
Stage 1: Week 1: SQL that answers questions
Joins, aggregation and row counts. Write 20 or more queries against a sample dataset.
- Typical effort
- 8–12 hours
- Outcome
- You can predict the row count of a join before running it.
Stage 2: Week 2: Window functions
Ranking, top N per group, LAG/LEAD and running totals.
- Typical effort
- 8–12 hours
- Outcome
- You can solve top-N-per-group and deduplication with window functions.
Stage 3: Week 3: Model the data
Facts, dimensions, grain and a star schema for one business process.
- Typical effort
- 8–12 hours
- Outcome
- You have a written star schema with a stated grain.
Stage 4: Week 4: Build a rerunnable load
Python loader with an upsert, a transaction, logging and tests.
- Typical effort
- 10–14 hours
- Outcome
- You have a loader whose test proves running it twice changes nothing.