Course · Distributed processing
Apache Spark
Understand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- Lessons
- 4
- Interview questions
- 9
- Projects & case studies
- 12
- Reading time
- ~1 h
About this course
Spark is a distributed processing engine. Your code builds a plan; Spark splits it into stages at shuffle boundaries and runs each stage as parallel tasks over partitions. Most tuning comes down to controlling how much data moves between executors and how evenly it is spread.
Start with partitions and shuffles, then move to join strategies and Adaptive Query Execution.
Your progress
Saved in this browser onlyPractise
- InterviewApache Spark interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Intermediate
Patterns used in production pipelines.
Advanced
Performance, internals and edge cases.
- Spark Partitions, Shuffles and Data SkewLearn how Spark splits data into partitions, why shuffles are expensive, how to recognise data skew and which fixes (AQE, broadcast, salting) apply.
- Spark Adaptive Query Execution and OptimizationHow Adaptive Query Execution re-plans Spark queries at run time: coalescing shuffle partitions, switching join strategies and splitting skewed partitions.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
- AdvancedKafka → Spark → Delta Lake Streaming PipelineAn application emits user events to Kafka. Build a streaming pipeline that lands them in Delta Lake within a minute, deduplicates replays, and produces per-minute aggregates that tolerate late events.
- IntermediateLarge-Scale Batch Processing PipelineProcess a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
- AdvancedReal-Time Analytics PipelineBuild a pipeline that turns a stream of order events into per-minute revenue and order counts by category, visible on a dashboard within a minute, and correct even when events arrive late.
System design case studies
- AdvancedDesign an A/B Testing Data PipelineDesign the data pipeline behind a company's experimentation platform: record which users saw which variant, join that to behavioural and business events, and produce daily, statistically sound results for hundreds of concurrent experiments.
- AdvancedDesign a Clickstream Data PlatformDesign a platform that collects every page view and click from a website and mobile apps and turns it into reliable product analytics such as sessions, funnels and retention.
- AdvancedDesign a Real-Time Streaming PlatformDesign a shared real-time streaming platform where hundreds of services publish domain events, platform users build stream-processing jobs on them, and the results reach the lakehouse, search, caches and alerting within seconds, reliably and with clear ownership.
- AdvancedDesign a Real-Time Analytics PipelineDesign a pipeline that turns application events into business metrics (orders per minute, revenue, conversion) visible on a dashboard within one minute of the events happening.
- IntermediateDesign a Scalable Batch Data PipelineDesign a pipeline that ingests daily order and customer extracts from several operational systems and produces reliable, analysis-ready tables for reporting by 07:00 each morning.
- AdvancedDesign a Lakehouse with Bronze, Silver and Gold LayersDesign a company-wide lakehouse in which many teams ingest batch files, database changes and event streams; data is refined through bronze, silver and gold layers; analysts query gold tables with SQL and data scientists train models from silver and gold, all on one governed copy of the data.
- AdvancedDesign a Streaming ETL Pipeline with Kafka and SparkDesign a streaming ETL pipeline that reads application events from Kafka, cleans, enriches and deduplicates them with Spark Structured Streaming, and lands them in lakehouse tables that analysts can query within a few minutes, without losing or double-counting events.
Resources
Cheat sheets
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- Cheat sheetApache Spark Interview Cheat SheetThe Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.
Related courses
- PySparkPySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.
- KafkaKafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.