Data Engineering · courses, interview preparation and projects
Learn Data Engineering. Prepare smarter. Get job ready.
Structured courses for every core technology, interview questions with direct answers, and projects you can explain, in one place.
- Courses
- 13
- Lessons
- 87
- Interview questions
- 100
- Projects
- 8
Your progress
Saved in this browser only, no account needed.Open a lesson or tick off an interview question and your progress appears here.
Continue where you left off:Courses
Pick a technology and start learning
Each course has ordered lessons, interview questions, projects and a cheat sheet.
- Languages & querySQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- Languages & queryPythonPython glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- Distributed processingPySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Distributed processingApache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- Streaming & orchestrationKafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- Streaming & orchestrationAirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- Data platformsDelta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.
- Data platformsData modelingData modeling and warehousing: star schemas and grain, fact and dimension design, SCDs, Data Vault and other methods, dbt, semantic layers and incremental models.
- Data platformsETL and ELTETL transforms data before loading it; ELT loads first and transforms inside the warehouse or lakehouse. Learn when each fits.
- Data platformsDatabricksDatabricks is a lakehouse platform built around Apache Spark and Delta Lake. Learn its workspace, compute, jobs and Unity Catalog governance model.
- Data platformsSnowflakeSnowflake is a cloud data warehouse that separates storage from compute. Learn virtual warehouses, micro-partitions, pruning, caching and cost control.
- CloudAWSAWS for Data Engineers: S3, IAM, Lambda, Glue, Athena, Redshift, EMR, Kinesis, Step Functions, Lake Formation and how they fit into one data platform.
- Interview foundationsDSADSA rounds in Data Engineering interviews are usually easy to medium problems on arrays, hashing, strings, two pointers and sliding windows, sometimes graphs or DP.
Roadmap
Data Engineer roadmap
From beginner foundations to job ready, in 13 ordered stages.
Interview
Three focused ways to prepare
- Technology interviewsInterview questionsDirect answers first, then reasoning, common mistakes and follow-ups.
- System designCase studiesRequirements, architecture, reliability and trade-offs for real data platforms.
- Company preparationCompany guidesPreparation that separates attributed information from practice questions.
Projects
Build something you can explain
- BeginnerCSV to Data Warehouse PipelineA small online shop exports orders as daily CSV files. Build a pipeline that loads them into a star schema so that sales can be reported reliably, even when files are resent or contain bad rows.
- IntermediateE-commerce Analytics Data PlatformAn online store has orders, customers, products and web sessions in separate systems. Build an ELT platform that models them into trusted marts for revenue, retention and product performance.
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
Featured guides
Start with these
- Career guideBeginnerHow to Explain a Data Engineering Project in an InterviewA repeatable structure for explaining a data engineering project in an interview: problem, design, trade-offs, failure handling and honest results.
- ArticleBeginnerStar, Snowflake and Galaxy Schemas and the GrainBuild a star schema for an online shop, declare the grain, snowflake a dimension, and query a galaxy of fact tables without double counting. Verified SQL.
- ArticleIntermediateKafka Topics, Partitions, Consumer Groups and Delivery SemanticsUnderstand how Kafka topics are split into partitions, how consumer groups share the work, and what at-most-once, at-least-once and exactly-once delivery really mean.
- ArticleIntermediatePySpark Window Functions: Ranking, Lag and Running TotalsUse PySpark window functions for top-N per group, previous-row comparisons and running totals, and avoid the single-partition trap of windows without partitionBy.
How we write
Editorial principles
Reviewed technical content
Claims about tools are checked against official documentation, and each page shows when it was last reviewed.
Version-aware
Where behaviour depends on a version or platform release, the page says which one it describes.
Evidence-labelled interviews
Company material is labelled as attributed, reported or representative practice. We never present invented questions as real ones.
Structured paths
Every lesson sits in a course with a clear previous and next step.