On this page
Tracks

System Design — Rate Limiter

Last reviewed 11 Sept 2026

Part of the system design series. See the framework and building blocks first if you haven’t.

1. Requirements

Functional

  • Limit how many requests a client (user, API key, or IP) can make in a time window.
  • Apply different rules per endpoint/tier (e.g. POST /login: 5/min, GET /search: 100/min).
  • Tell the client it was throttled, ideally with when it can retry.

Non-functional

  • Low added latency (single-digit ms) — it sits in front of every request.
  • Accurate enough, not perfectly exact — over-blocking a legitimate burst by a few requests is fine; a limiter that stalls the whole API is not.
  • Must work across many app servers behind a load balancer (a per-process counter is useless the moment you scale horizontally).
  • Must not become the outage — if the limiter’s own dependency dies, the API should degrade, not disappear.

2. Where it lives

flowchart LR
C[Client] --> GW["API Gateway / Edge<br/>(Cloudflare, Kong, Envoy)"]
GW --> LB[Load Balancer]
LB --> S1[App Server]
LB --> S2[App Server]
S1 --> RD[(Shared counter store<br/>Redis)]
S2 --> RD
Placement options
  • Edge / gateway — stops abusive traffic before it costs you anything (bandwidth, app CPU). This is where Cloudflare enforces its network-wide limits.
  • Middleware in the app — needed for business-logic-aware limits (per-user plan, per-API-key tier) that the edge doesn’t know about.
  • Most production systems run both: a coarse edge limiter for DoS/abuse, a finer app-level limiter for business rules.

3. Algorithms

AlgorithmMemoryAccuracyBurst behaviorUsed by
Fixed windowO(1)Weak — up to 2x burst at window edgeAllows a burst at the boundarySimple internal APIs
Sliding logO(requests)ExactNoneLow-volume, high-precision needs
Sliding window counterO(1)~99.99% accurate in practiceSmooths the boundary burstCloudflare (default at their scale)
Token bucketO(1)Exact, allows configurable burstsBursty by design (spend saved tokens)Stripe (client-side limiter)
GCRA (leaky-bucket variant)O(1), one timestamp per keyExactSmooth, no burstStripe, Vimeo, Doorman-style systems

GCRA in one paragraph: instead of counting requests, track a single per-key timestamp — the “theoretical arrival time” (TAT) of the next request if traffic were perfectly smooth. Each request compares now to TAT: if now >= TAT - burst_allowance, allow it and push TAT forward by the cost of one request; otherwise reject. One integer per key, one comparison, no buckets to reset. This is why it’s the production choice at companies that care about both accuracy and memory — Stripe and Vimeo both run GCRA-family limiters.

4. Distributed enforcement

A single Redis (or Redis Cluster) instance as shared state is the standard production pattern:

  • INCR + EXPIRE, or a Lua script that does check-and-increment atomically in one round trip — this is the part people get wrong: a naive GET-then-SET from the app has a race condition under concurrent requests from the same key. The Lua script runs entirely inside Redis, so it’s atomic without a distributed lock.
  • Sub-5ms latency at the scale most systems need.
  • For GCRA specifically, the Redis script only needs to read/write one TAT value per key — this is what redis-gcra-style libraries do.
sequenceDiagram
participant Client
participant App as App Server
participant Redis
Client->>App: Request (key = user:123)
App->>Redis: EVAL rate_limit.lua(key, limit, window)
Redis-->>App: allow / deny + retry_after
alt allowed
  App-->>Client: 200 + X-RateLimit-Remaining
else denied
  App-->>Client: 429 + Retry-After
end
Request path with a Redis-backed limiter

Return the IETF RateLimit-* headers (RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset) plus Retry-After on a 429 — this is what Stripe, GitHub, and most public APIs standardized on, so clients can back off correctly instead of hammering you again immediately.

5. What happens when Redis goes down

This is the question interviewers love, because most candidates haven’t thought about it.

  • Fail closed (reject everything if Redis is unreachable): you just turned your own rate limiter into a self-inflicted outage.
  • Fail open (allow everything if Redis is unreachable): safe for availability, but the moment Redis died because of load, you just removed the one thing protecting the DB from that load — right when it’s needed most.
  • What real systems do: catch the Redis error at every layer so a limiter bug or outage never 5xx’s the actual request — default to fail-open, but pair it with a local, in-process fallback limiter (a simple in-memory token bucket per app server) so there’s still some ceiling during the outage, just a coarser, per-instance one instead of a precise global one.
  • Google’s Doorman takes a more elegant approach for its use case: each instance enforces the last-known drop ratio pushed by a central controller, and keeps enforcing that stale ratio if the controller is unreachable — never fully open, never fully closed, degrading gracefully instead of binary-flipping.

6. Scaling further

  • Redis becomes the bottleneck → shard by key (hash(user_id) mod N across a Redis Cluster) so no single node holds every key.
  • Cross-region traffic → don’t ship every request to one global Redis; run a limiter per region with a coarser global budget synced asynchronously (this is close to what Doorman does — local enforcement, async coordination).
  • Multi-tenant, wildly different limits per plan → keep the algorithm generic, store (limit, window) per key’s tier in a small config cache (or embed it in the JWT) so you’re not doing a config lookup on every request.

Interview follow-ups

  • “Fixed window vs sliding window — walk me through the boundary problem.” — Use the 99+99 example above with concrete numbers.
  • “Where would you put this — client, gateway, or app server? Why not just one?” — Edge for cheap abuse rejection, app layer for business-aware limits; most real systems run both.
  • “Redis just went down. What happens to your API?” — Fail-open + local in-memory fallback; explain the trade-off, don’t just pick one.
  • “How do you avoid a race condition when two requests for the same key hit different app servers at the same instant?” — Atomicity has to live in the shared store (Lua script / INCR), not in app code.
  • “10x traffic overnight. What’s the first thing that breaks?” — The shared Redis instance’s connection count and CPU before the algorithm itself; shard it.
  • “How would a client know it got rate-limited and when to retry, without guessing?” — RateLimit-* headers + Retry-After.

Sources: Design a Distributed Rate Limiter — Hello Interview · Rate Limiting, Cells, and GCRA — brandur.org · Scaling your API with rate limiters — Stripe · Doorman design doc — YouTube/Google · Fail-Open vs Fail-Closed Middleware