Home / Software y Cloud / Inside Kafka: the event log that keeps streaming alive without losing a single message

Inside Kafka: the event log that keeps streaming alive without losing a single message

Ilustracion de un log de eventos distribuido de Apache Kafka

When an online shopping cart, an industrial sensor or your click history leaves a trace in almost any modern architecture, the message almost always ends up in the same place: an append-only log. That log is called Kafka when we talk about the streaming platform moving millions of messages per second, and understanding how it records (and never erases) is the key to why the internet has become able to enqueue almost everything.

This article goes into the real mechanics of Kafka’s model: partitions, replicas, offsets, consumer groups and the compacted log that gives every message a permanent address.

A message is not forwarded: it is read and you move on

Kafka’s core idea is almost offensively simple: instead of “sending” messages to recipients the way a classic queue does (the point-to-point pattern of RabbitMQ or a Java Message Service queue), Kafka records each message at the end of a continuous ledger. The producer only writes; the consumer only reads.

That ledger is called an append-only log: data is only added at the end, never inserted mid-way or rewritten. It is the same idea databases already use with their WAL (write-ahead log, the pre-write record that guarantees no transaction is lost before being stored): writing sequentially is the cheapest thing a disk can do.

Partitions: the parallelism that brings scale

A single log on a single disk does not scale. That is why Kafka splits every topic (the category of messages, like “orders” or “visits”) into several partitions. Each partition is an independent, ordered log; Kafka’s parallelism comes from being able to read and write many partitions at once across different servers.

Ordering is guaranteed within a partition, not across partitions. If you need the events of one client to be processed in order, you make sure on your own that they all land in the same partition (for example, by hashing the client key). It is one of the design decisions that sets Kafka apart from queues with global ordering.

Replicas and the leader: surviving a dead disk

Data recorded only in a local log is lost if that machine dies. Kafka’s fault tolerance rests on replicas: each partition is copied across several cluster nodes and one of them acts as the leader. Writers and readers talk to the leader; the rest are followers that replicate the data.

Under the hood it is the same quorum strategy of distributed systems: accepting and confirming a write only requires the majority of replicas, because Kafka uses the Raft consensus protocol to elect leaders and keep the cluster state. If the leader fails, the followers elect another and the write is not lost as long as it was replicated on the majority.

The offset: where you left off last time

Here is the magic a classic queue cannot give you. Every message inside a partition has an offset: a sequence number that acts as its exact address. Since nothing is erased when read, a consumer can read message 100, crash, and resume at message 101 with nothing lost: it only needs to remember its last offset.

That allows reprocessing an entire batch by resetting the offset, or letting several consumers read the same partition at different times without interfering. When a consumer group distributes the partitions among its members, each consumer takes responsibility for a subset and stores its offsets like bookmarks: that is the mechanism behind “each message is delivered to a single member of the group”, Kafka’s counterpart to load balancing.

The twist: log compaction

An infinite log would grow without bound. Kafka keeps messages for a configurable retention period, but it also has a special mode called log compaction: instead of deleting the old, it keeps only the latest version of each key. If you write “price of product A = 10” and then “price of product A = 15”, compaction leaves only the 15, because the only thing that matters to rebuild the final state is the last value.

That turns a topic into an event table with which you can reconstruct the exact state of a system at any point in the past, and it is the pattern known as event sourcing. The log acts simultaneously as an event database and a message queue, which is why it earned the reputation of “backbone of event streaming”.

Why everyone talks about Kafka

Kafka was not the first to use an append-only log, but it standardized the model at internet scale and is now practically the glue of data architectures: it captures every click and event, hands them out to processors (often with stream-processing tools such as Kafka Streams or Flink) and feeds everything from analytics dashboards to model training.

Understanding Kafka is understanding the difference between “sending a message that another party must collect” and “recording a fact that anyone can read whenever they want, in the order it happened”. That difference —the immutable log instead of the volatile queue— is what keeps millions of messages per second from ever being lost.