Menu

Kafka course · Lesson 1 of 3

Kafka and Real-Time Data Engineering

Real-time data engineering with Kafka: logs, partitions and consumer groups, delivery semantics, schemas, stream processing, CDC and operating streaming pipelines.

  • Intermediate
  • Pillar guide
  • 2 min read
  • Updated Oct 2026
On this page
  1. 1. Kafka’s model
  2. 2. Log or queue?
  3. 3. Delivery semantics
  4. 4. Schemas and contracts
  5. 5. Stream processing
  6. 6. Change data capture
  7. 7. Operations
  8. Build it

Real-time data engineering means processing events continuously instead of in scheduled batches. Kafka is the most common backbone; this guide connects the concepts from transport to processing to operations.

1. Kafka’s model

A topic is a partitioned, replicated, append-only log. Partitions provide ordering and parallelism; keys route related events to the same partition; consumer groups divide partitions among consumers.

Read: Topics, partitions and consumer groups · Practise: Partitions and consumer groups

2. Log or queue?

Logs retain events for replay by many consumers; queues distribute units of work. Most data pipelines want the log model.

Read: Kafka vs message queues

3. Delivery semantics

At-least-once is the practical default; make processing and sinks idempotent. Kafka’s exactly-once covers read-process-write within Kafka.

Practise: At-least-once delivery

4. Schemas and contracts

Use a schema registry with compatibility rules so producers cannot break consumers.

Read: Schema evolution patterns

5. Stream processing

Aggregate by event time, use watermarks for late data, deduplicate by event id, checkpoint state and offsets.

Read: Batch vs streaming · Case study: Real-time analytics pipeline

6. Change data capture

Stream database changes from the transaction log through Kafka and apply them with ordered, idempotent merges.

Read: CDC patterns and failure modes · Case study: CDC platform

7. Operations

Monitor consumer lag, under-replicated partitions and end-to-end latency; size partitions for peak throughput and consumer parallelism; set retention longer than your longest outage.

Read: Kafka ingestion system design

Build it

Start with the Kafka → Spark → Delta streaming project, then the real-time analytics project. Revise with the Kafka cheat sheet.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Concepts apply to Apache Kafka 3.x and later

Progress is saved in this browser only. No account needed.

Search
Filter by type