ProjectsProject 4 of 8
Intermediate project · Project 4 of 8
Large-Scale Batch Processing Pipeline
Process a large public dataset (several gigabytes or more) with PySpark into partitioned, query-ready tables, and document how you found and fixed the main performance bottleneck.
Requirements
- Read a dataset of at least several gigabytes
- Clean, join with a lookup table and aggregate
- Write partitioned Parquet or Delta output by date
- Rerun any date without duplicating output
- Record a before-and-after performance analysis from the Spark UI
Technology stack
PySpark, Parquet or Delta Lake, Local Spark or a small cloud cluster
Dataset
Use a large public dataset with a clear licence, such as public trip records or open web-analytics samples. Record the source and licence in your README.
Business context
Interviewers for Spark roles want evidence that you have worked with data too big for a laptop’s memory and know how to find bottlenecks. This project produces exactly that evidence: a working job and a written performance analysis.
Architecture
- Raw files in a
raw/folder or bucket. - PySpark job parameterised by processing date.
- Broadcast join with a small lookup table.
- Aggregation to the reporting grain.
- Partitioned output with overwrite for the processed date.
Write the performance analysis as you go: screenshots of the stage timeline, the task-duration spread, shuffle sizes and what changed after your fix. Report the timings you actually observed on your hardware.
Progress is saved in this browser only. No account needed.