Menu

Advanced project · Project 6 of 8

Change Data Capture Pipeline

Replicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.

  • Advanced
  • PostgreSQL · A log-based CDC connector (for example Debezium) · Kafka · Spark Structured Streaming · Delta Lake · Docker Compose
  • 2 min read
  • Updated Oct 2026

Requirements

  • Run PostgreSQL with logical replication enabled
  • Capture changes with a log-based CDC connector into Kafka
  • Apply changes to a Delta table with MERGE, ordered by source log position
  • Handle deletes
  • Take an initial snapshot without missing concurrent changes

Technology stack

PostgreSQL, A log-based CDC connector (for example Debezium), Kafka, Spark Structured Streaming, Delta Lake, Docker Compose

Dataset

Create your own orders table and a small script that inserts, updates and deletes rows continuously.

Business context

Analysts need current operational data, but querying the production database directly risks slowing the application. CDC keeps an analytical copy up to date from the database’s own change log, without extra query load.

Architecture

  1. PostgreSQL writes changes to its write-ahead log.
  2. The CDC connector reads the log and publishes change events to Kafka.
  3. A streaming job reads micro-batches and deduplicates to the latest change per key.
  4. MERGE applies inserts, updates and deletes to the Delta table.
  5. Analysts query the Delta table.
Changes flow from the database log to the lakehouse with per-key ordering.

The design follows the CDC system design case study; read it before building.

By Data Career Hub Editorial · Last reviewed Oct 2026

Progress is saved in this browser only. No account needed.

Search
Filter by type