Menu

Kafka course · Lesson 3 of 3

Kafka vs Queue-Based Messaging for Data Pipelines

When to use a distributed log like Kafka and when a traditional message queue fits better: retention, replay, ordering, fan-out, per-message acknowledgement and operations.

  • Intermediate
  • 2 min read
  • Updated Oct 2026
On this page
  1. Choose a log when
  2. Choose a queue when
  3. In data engineering
  4. Common mistakes
  5. Key takeaway

Both move messages between systems. They are built around different models, and that decides which fits a data pipeline.

Distributed log (Kafka) Message queue (typical)
Model Append-only log; consumers track their own position Messages delivered to a consumer and removed after acknowledgement
Retention Time- or size-based, independent of consumption Usually until consumed (or expired)
Replay Yes: reset the offset and read again Generally no, once acknowledged
Many independent consumers Natural: each consumer group reads everything Needs fan-out (topics/exchanges) configured per subscriber
Ordering Per partition Varies; often per queue, weakened by multiple consumers
Work distribution By partition Per message, flexible
Typical strengths High-throughput event streams, CDC, replayable pipelines Task queues, per-message work, request decoupling

Choose a log when

  • Several systems need the same events (lakehouse, search, alerts).
  • You need to replay history after a bug or to bootstrap a new consumer.
  • Throughput is high and events are part of an ongoing stream.
  • Per-key ordering matters (all changes for one order in sequence).

Choose a queue when

  • Each message is a unit of work that one worker should do once (resize an image, send an email).
  • You want per-message acknowledgement, retries and dead-lettering without managing offsets.
  • Volumes are moderate and replay is not needed.

In data engineering

Most pipeline use cases (event streaming, CDC, feeding a lakehouse) fit the log model because replay and multiple consumers are valuable. Task orchestration between pipeline steps is usually handled by an orchestrator, not by either.

Common mistakes

  1. Using a queue where consumers later need to replay history.
  2. Using Kafka for a small task queue and taking on its operational overhead.
  3. Expecting global ordering from either.

Key takeaway

Logs retain and replay events for many consumers; queues distribute units of work and forget them. Pick by replay, fan-out and ordering needs.

By Data Career Hub Editorial · Last reviewed Oct 2026 · Conceptual; individual queue products differ, and Kafka has added queue-like consumption features in recent releases

Progress is saved in this browser only. No account needed.

Search
Filter by type