Menu

Beginner project · Project 1 of 8

CSV to Data Warehouse Pipeline

A small online shop exports orders as daily CSV files. Build a pipeline that loads them into a star schema so that sales can be reported reliably, even when files are resent or contain bad rows.

  • Beginner
  • Python · SQLite or PostgreSQL · SQL · pytest
  • 2 min read
  • Updated Oct 2026

Requirements

  • Load daily order CSV files into a local database
  • Model the data as a fact table and at least two dimensions
  • Reruns must not create duplicates
  • Bad rows are logged and counted, not silently dropped
  • Automated tests prove idempotency

Technology stack

Python, SQLite or PostgreSQL, SQL, pytest

Dataset

Generate your own CSV files with a small script, or use any public sample retail dataset whose licence allows reuse.

Business context

Reports built from hand-edited spreadsheets are slow and error-prone. A small, reliable pipeline with a clear model gives the shop consistent numbers every morning and is a good first project because every Data Engineering concept appears in miniature: modelling, idempotency, quality and testing.

Architecture

  1. Daily CSV files land in a data/ folder.
  2. A Python loader writes them to a staging table with upserts and an audit record.
  3. SQL builds dim_customer, dim_product and fact_order_line.
  4. Quality checks run; failures stop the report refresh.
A single-machine batch pipeline with staging, modelling and checks.

Start from the idempotent CSV loader tutorial, then add the modelling layer described in the star schema guide.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type