data-engineering interview questionsQuestion 6 of 6
Data Engineering interview question · Question 6 of 6
How would you design a CDC pipeline?
Short answer
I would use log-based change data capture: a connector reads the database's transaction log and publishes one event per row change, with the operation, row values and log position, to Kafka topics keyed by primary key so each row's changes stay in order. A consumer applies micro-batches to the target table with MERGE, keeping only the latest position per key and ignoring events older than what is stored, which makes replays harmless. I also plan the initial snapshot, delete handling, schema changes and monitoring of lag.
Detailed explanation
Walk through the design in this order:
- Capture: log-based CDC (reads the transaction log, captures deletes, adds no query load) rather than polling
updated_at(misses deletes and intermediate changes). - Transport: Kafka topic per table, keyed by primary key, so ordering holds per row.
- Apply: micro-batch
MERGEinto the target. Deduplicate to the latest log position per key; update only when the incoming position is newer. - Bootstrap: consistent snapshot plus the log position at snapshot time; stream from that position.
- Deletes: hard delete, or a soft-delete flag if history is needed.
- Schema changes: additive changes flow through; breaking changes pause and alert.
- Operations: monitor connector lag, consumer lag and end-to-end latency; keep Kafka retention longer than the longest expected outage.
See the full CDC system design case study.
Common mistakes
- Applying events in arrival order without comparing log positions.
- Forgetting deletes.
- No plan for the initial load.
Progress is saved in this browser only. No account needed.