Course · Distributed processing
PySpark
PySpark is the Python API for Apache Spark. Learn DataFrames, joins, window functions and how partitions and shuffles decide performance.
- Lessons
- 6
- Interview questions
- 7
- Projects & case studies
- 2
- Reading time
- ~1 h
About this course
PySpark lets you write distributed data processing in Python. DataFrames look familiar if you know SQL or pandas, but they are executed lazily across a cluster, so how data is partitioned and moved decides whether a job takes minutes or hours.
Learn DataFrames and joins first, then window functions, then the execution model (jobs, stages, tasks) and how to read the Spark UI.
Your progress
Saved in this browser onlyPractise
- InterviewPySpark interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Beginner
Core concepts you will use every day.
- PySpark DataFrames and SchemasCreate PySpark DataFrames with explicit schemas, understand nullability and type inference, and choose how malformed records are handled when reading files.
- PySpark Transformations vs ActionsUnderstand lazy evaluation in PySpark: transformations build a plan, actions run it, and narrow versus wide transformations decide where Spark shuffles data.
Intermediate
Patterns used in production pipelines.
- PySpark Joins and Join StrategyWrite correct PySpark joins (inner, left, anti, semi), avoid duplicate-column and fan-out bugs, and understand when Spark broadcasts or sort-merges a join.
- PySpark Window Functions: Ranking, Lag and Running TotalsUse PySpark window functions for top-N per group, previous-row comparisons and running totals, and avoid the single-partition trap of windows without partitionBy.
- PySpark UDFs and Safer AlternativesWhen Python UDFs are slow or risky in PySpark, how built-in functions and pandas UDFs compare, and what Arrow-optimised UDFs in Spark 4 change.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- IntermediateLarge-Scale Batch Processing PipelineProcess a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
- IntermediateS3 → PySpark → Snowflake Data PipelineRaw event files arrive in object storage every day. Clean and aggregate them with PySpark and load curated tables into a cloud warehouse for analysts, with each day reprocessable on demand.
Resources
Cheat sheets
- Cheat sheetPySpark Cheat SheetA quick PySpark reference: reading and writing data, column expressions, joins, aggregations, window functions and the settings that matter for performance.
- Cheat sheetApache Spark Interview Cheat SheetThe Spark concepts interviewers ask about most, in one page: lazy evaluation, stages and shuffles, joins, partitions, skew, AQE, caching and Spark 4 defaults.
Related courses
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- SQLSQL is the core language of data work: querying, transforming and modelling data in warehouses, lakehouses and Spark. Start here before any other tool.
- PythonPython glues pipelines together: ingestion, validation, orchestration and PySpark jobs. Focus on functions, generators, error handling and testable code.
- Delta LakeDelta Lake adds ACID transactions, schema enforcement and time travel to files in a data lake, which is the foundation of the lakehouse pattern.