ProjectsProject 2 of 8
Intermediate project · Project 2 of 8
S3 → PySpark → Snowflake Data Pipeline
Raw event files arrive in object storage every day. Clean and aggregate them with PySpark and load curated tables into a cloud warehouse for analysts, with each day reprocessable on demand.
Requirements
- Read one day of raw JSON or CSV files from object storage
- Clean, deduplicate and aggregate with PySpark
- Write curated Parquet partitioned by date
- Load the curated data into a warehouse table idempotently
- Document cost and access controls
Technology stack
PySpark, Amazon S3 (or any object storage), Snowflake (or another cloud warehouse), Airflow (optional)
Dataset
Use a public event or trip dataset whose licence allows reuse, or generate synthetic events.
Business context
Most production batch pipelines look like this: files land in object storage, a distributed job cleans them, and a warehouse serves analysts. Building it end to end teaches the boundaries between storage, processing and serving.
Architecture
- Raw files land in
s3://bucket/raw/dt=YYYY-MM-DD/. - PySpark reads one date, cleans and aggregates.
- Curated Parquet is written to
curated/dt=YYYY-MM-DD/with overwrite. - The warehouse table for that date is replaced in a transaction.
Use least-privilege credentials for the job, never hard-code keys, and keep secrets in your platform’s secret manager.
Progress is saved in this browser only. No account needed.