Course · Streaming & orchestration
Kafka
Kafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.
- Lessons
- 3
- Interview questions
- 3
- Projects & case studies
- 10
- Reading time
- ~1 h
About this course
Kafka stores events in partitioned, replicated logs that many consumers can read independently. It decouples producers from consumers and is the backbone of many streaming and change-data-capture designs.
Learn how partitions and consumer groups relate before you tune anything. Most Kafka interview questions come back to those two ideas.
Your progress
Saved in this browser onlyPractise
- InterviewKafka interview questionsThe full list with difficulty, type and a box to tick off each one.
- Cheat sheetKafka Cheat SheetA quick Kafka reference: topics, partitions, keys, replication, consumer groups, offsets, delivery semantics and the CLI commands for inspecting a cluster.
- InterviewAll interview questionsEvery question across all topics in one filterable list.
Course structure
Lessons
Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.
Start here
The complete overview of the course in one read.
Intermediate
Patterns used in production pipelines.
- Kafka Topics, Partitions, Consumer Groups and Delivery SemanticsUnderstand how Kafka topics are split into partitions, how consumer groups share the work, and what at-most-once, at-least-once and exactly-once delivery really mean.
- Kafka vs Queue-Based Messaging for Data PipelinesWhen to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.
Projects and case studies
Apply what you learned and prepare material to discuss in interviews.
Projects
- AdvancedChange Data Capture PipelineReplicate an operational PostgreSQL table into a lakehouse table within minutes, including updates and deletes, so analysts query current data without touching the production database.
- AdvancedFraud Detection Data PipelineBuild the data side of a fraud-detection system: compute per-card behavioural features from a transaction stream, flag suspicious transactions with transparent rules, and maintain a feature table that a model could use.
- AdvancedKafka → Spark → Delta Lake Streaming PipelineAn application emits user events to Kafka. Build a streaming pipeline that lands them in Delta Lake within a minute, deduplicates replays, and produces per-minute aggregates that tolerate late events.
- AdvancedReal-Time Analytics PipelineBuild a pipeline that turns a stream of order events into per-minute revenue and order counts by category, visible on a dashboard within a minute, and correct even when events arrive late.
System design case studies
- AdvancedDesign a CDC Pipeline from an OLTP Database to the WarehouseReplicate inserts, updates and deletes from a production PostgreSQL (or MySQL) database into the analytics warehouse within minutes, keeping both a current-state copy and a change history, without adding query load to the source or losing a single change.
- AdvancedDesign a Clickstream Data PlatformDesign a platform that collects every page view and click from a website and mobile apps and turns it into reliable product analytics such as sessions, funnels and retention.
- AdvancedDesign an Event-Driven ArchitectureAn e-commerce company's services call each other synchronously, so one slow service stalls checkout and analytics depends on nightly database dumps. Design an event-driven architecture in which services publish business events reliably, other services react to them independently, and the same events feed analytics, without losing or double-applying anything.
- AdvancedDesign a Real-Time Streaming PlatformDesign a shared real-time streaming platform where hundreds of services publish domain events, platform users build stream-processing jobs on them, and the results reach the lakehouse, search, caches and alerting within seconds, reliably and with clear ownership.
- AdvancedDesign a Real-Time Analytics PipelineDesign a pipeline that turns application events into business metrics (orders per minute, revenue, conversion) visible on a dashboard within one minute of the events happening.
- AdvancedDesign a Streaming ETL Pipeline with Kafka and SparkDesign a streaming ETL pipeline that reads application events from Kafka, cleans, enriches and deduplicates them with Spark Structured Streaming, and lands them in lakehouse tables that analysts can query within a few minutes, without losing or double-counting events.
Resources
Cheat sheets
Related courses
- Apache SparkUnderstand how Spark turns your code into jobs, stages and tasks, and why partitions, shuffles and data skew drive performance.
- AirflowAirflow schedules and orchestrates pipelines as DAGs. Learn scheduling, task dependencies, retries and idempotent task design.
- Data modelingData modeling and warehousing: star schemas and grain, fact and dimension design, SCDs, Data Vault and other methods, dbt, semantic layers and incremental models.