Menu

Course · Streaming & orchestration

Kafka

Kafka is a distributed log used for streaming data. Learn topics, partitions, consumer groups and delivery semantics before building streaming pipelines.

Lessons
3
Interview questions
3
Projects & case studies
10
Reading time
~1 h

About this course

Kafka stores events in partitioned, replicated logs that many consumers can read independently. It decouples producers from consumers and is the backbone of many streaming and change-data-capture designs.

Learn how partitions and consumer groups relate before you tune anything. Most Kafka interview questions come back to those two ideas.

Your progress

Saved in this browser only

Practise

Course structure

Lessons

Work through the lessons in order. Completed lessons show a tick; lessons you have opened are outlined.

Start here

The complete overview of the course in one read.

  1. Kafka and Real-Time Data EngineeringReal-time data engineering with Kafka: logs, partitions and consumer groups, delivery semantics, schemas, stream processing, CDC and operating streaming pipelines.Intermediate2 min

Intermediate

Patterns used in production pipelines.

  1. Kafka Topics, Partitions, Consumer Groups and Delivery SemanticsUnderstand how Kafka topics are split into partitions, how consumer groups share the work, and what at-most-once, at-least-once and exactly-once delivery really mean.Intermediate4 min
  2. Kafka vs Queue-Based Messaging for Data PipelinesWhen to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.Intermediate2 min

Projects and case studies

Apply what you learned and prepare material to discuss in interviews.

Projects

System design case studies

Resources

Cheat sheets

Related courses

Plan your learning

Search
Filter by type