Menu

Intermediate project · Project 2 of 8

S3 → PySpark → Snowflake Data Pipeline

Raw event files arrive in object storage every day. Clean and aggregate them with PySpark and load curated tables into a cloud warehouse for analysts, with each day reprocessable on demand.

  • Intermediate
  • PySpark · Amazon S3 (or any object storage) · Snowflake (or another cloud warehouse) · Airflow (optional)
  • 2 min read
  • Updated Oct 2026

Requirements

  • Read one day of raw JSON or CSV files from object storage
  • Clean, deduplicate and aggregate with PySpark
  • Write curated Parquet partitioned by date
  • Load the curated data into a warehouse table idempotently
  • Document cost and access controls

Technology stack

PySpark, Amazon S3 (or any object storage), Snowflake (or another cloud warehouse), Airflow (optional)

Dataset

Use a public event or trip dataset whose licence allows reuse, or generate synthetic events.

Business context

Most production batch pipelines look like this: files land in object storage, a distributed job cleans them, and a warehouse serves analysts. Building it end to end teaches the boundaries between storage, processing and serving.

Architecture

  1. Raw files land in s3://bucket/raw/dt=YYYY-MM-DD/.
  2. PySpark reads one date, cleans and aggregates.
  3. Curated Parquet is written to curated/dt=YYYY-MM-DD/ with overwrite.
  4. The warehouse table for that date is replaced in a transaction.
Every stage is keyed by processing date so any day can be rerun independently.

Use least-privilege credentials for the job, never hard-code keys, and keep secrets in your platform’s secret manager.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type