Home / Software y Cloud / 429 Too Many Requests: the bouncer that decides how much you are allowed to ask

429 Too Many Requests: the bouncer that decides how much you are allowed to ask

When you fire thousands of requests at an API in a single second, most people expect the server to answer every single one. A server that always answers, with no brakes, is designed to collapse. Before that happens, something has to say no. That something is the rate limiter, and behind its most famous reply — 429 Too Many Requests — lie three algorithms worth knowing.

The problem underneath

An API exposes a finite service: every request consumes CPU, memory, database access or calls to other systems. If one client (or an attacker) floods the service with traffic, legitimate requests run out of resources. That is why rate limiters are not a matter of courtesy: they are a capacity contract.

The goal is to allow a steady pace with controlled spikes. Two metrics rule everything: the rate (how many requests per second are admitted on average) and the burst size (the momentary peak that is tolerated without rejection). The algorithms differ in how they balance those two figures.

2. Token bucket, the generous one

The token bucket models the limit as a reservoir with a fixed capacity capacity that is filled with tokens at a constant rate rate (tokens per second). Each request spends one token; if the bucket is empty, the request is rejected.

The key is that tokens accumulate. If nothing happens for a while, the bucket reaches its top and the server can absorb a peak of up to capacity instantaneous requests. It is the natural choice when you want to allow short bursts without punishing calm.

In pseudocode, the logic is surprisingly small:

def allow():
    now = time()
    tokens = min(capacity, tokens + (now - last) * rate)
    last = now
    if tokens >= 1:
        tokens -= 1
        return True
    return False

The token accumulation (now - last) * rate compensates for idle time, and min keeps the bucket from growing without limit.

3. Leaky bucket, the one that evens the pace

The leaky bucket thinks the other way: requests enter at the top and are served at a fixed rate from the bottom, like a bucket with a hole. If the bucket fills up, the extra requests are discarded or queued.

Unlike the token bucket, the output here is constant: there are no bursts. It is perfect when the resource downstream (a disk write rate, a billing ceiling) must never exceed a certain value. The downside is rigidity: a calm user with a sudden peak gets rejected all the same.

There is a direct relationship between both: a token bucket with rate = r and capacity = r behaves almost like a leaky bucket draining at a fixed pace. The difference is the fate of unused tokens: accumulate (token) or overflow (leaky).

4. Fixed and sliding windows: the fight against the edge

Window algorithms count requests in a time interval. A fixed window divides time into slices (say 60 seconds) and counts requests per slice. It is cheap and simple, but has a well-known flaw at the clock edge:

If the limit is 100 requests per minute, a client can fire 100 in the last 5 seconds of minute 1 and another 100 in the first 5 of minute 2: 200 requests in 10 seconds, double the capacity. The token bucket and the sliding window fix the offset.

Real implementations use a sliding log window: each request’s timestamp is stored and only those in the current period are counted. Memory grows with traffic, so in production a sliding window approximation is preferred, which averages the request between the previous and the current window: accurate at the boundary and cheap on CPU.

5. GCRA: what Redis actually uses

Most distributed limiters do not hand-code these algorithms: they use GCRA (Generic Cell Rate Algorithm), a telecom-grade algorithm that combines token bucket and time reservation. Its advantage is that it stores only two numbers per key: the instant the next request should be allowed and the theoretical emission time.

In an environment with several servers behind a load balancer, the state must be shared. The standard solution is Redis with an atomic Lua script: the counter (INCR with EXPIRE) or the GCRA state is read and modified in a single atomic operation, avoiding race conditions across nodes. Some providers already expose this mechanism as a managed service.

6. The 429 and the Retry-After header

When the limit is exceeded, the server has to say so with an explicit code. The status 429 Too Many Requests was defined in RFC 6585, and the server can add the Retry-After header telling the client how many seconds (or a date) to wait before retrying.

In practice more escape headers are added: X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset, so the client can pace itself before being rejected. A well-behaved client respects Retry-After and applies exponential backoff: after a 429 it waits longer and longer between attempts, instead of hammering the server and making things worse.

The takeaway

Behind every 429 there is a design decision: do I accumulate tokens to reward calm (token bucket)? Do I even out the pace no matter what (leaky bucket)? Do I count in fixed slices and accept the edge problem (fixed window)? The choice depends on the resource being protected, not on whim.

And when the API is spread across many machines, the real difficulty is not the algorithm but shared state: that is why Redis and an atomic Lua script have become the de facto standard. The next 429 you see will no longer be just an error: it will be a conversation between your client and a doorman who knows exactly how much patience is left.