Rate Limiting
A bot with a broken retry loop hammers Adda's search endpoint 50,000 times a second and drowns everyone else. Nabila draws the line: rate limiting caps how much any one client can ask for.
The problem
Adda's public API hums along at a few thousand requests per second. Then one integration — a third-party app built on Adda's API — ships a buggy retry loop with no backoff and starts hammering the search endpoint 50,000 times a second. Shuvo's database saturates, latency for everyone climbs, and requests start timing out. Fahim can't load his feed again.
The infuriating part: one client's mistake is now everyone's outage. Legitimate users can't log in because one script is monopolizing the capacity they all share. Nabila doesn't want to ban the client; she just wants to say "you get your fair slice, and no more." Adda needs a line that caps how much anyone can ask for in a given window.
A first attempt
Count requests per client and reject anything over the limit. The simplest version: a fixed window — "100 requests per minute." Keep a counter per client that resets at the top of each minute.
counter[client]++ each request
if counter[client] > 100: reject (429)
reset all counters at :00 of every minuteIt mostly works, but has a nasty edge. The window is a hard boundary, so a client can send 100 requests at 12:00:59 and another 100 at 12:01:00 — 200 requests in one second, right across the reset. Bursts at the seam sail straight through the limit you thought you set.
The insight
Stop thinking in hard-resetting buckets and think in rate. The most flexible model is the
token bucket: a bucket holds up to B tokens and refills at R tokens per second. Each
request must take one token; if the bucket is empty, the request is rejected.
This captures what Nabila actually wants. R sets the sustained rate; B sets how big a burst
she'll tolerate. A client that's been quiet accumulates tokens and can spike briefly (good UX
for a bursty mobile app), but can't exceed R on average over time. One dial for steady rate,
one for burst.
Rate limit on rate, plus a burst allowance
Token bucket = refill rate R (long-run cap) + bucket size B (instantaneous burst). It
smooths traffic without the reset-boundary loophole, which is why it's the default choice for
most APIs.
How it works
Give each client a bucket
Keyed by API key, user ID, or IP. The bucket has a capacity B and starts full. Store two
numbers: current token count and the timestamp of the last refill.
Refill lazily on each request
Don't run a timer. When a request arrives, add (now − last_refill) × R tokens (capped at B)
and update the timestamp. This computes the refill on demand — no background job.
Take a token or reject
If tokens ≥ 1, subtract one and allow the request. If not, reject with HTTP 429 Too Many
Requests and a Retry-After header telling the client when a token will be available.
Do it centrally for a fleet
With Adda's many gateway servers, per-server counters let a client multiply their limit by the server count. Keep the bucket in a shared store (Redis) so all servers see one count — usually an atomic Lua script so the read-modify-write can't race.
Concrete numbers
Set R = 100/s and B = 200. A steady client runs all day at 100 req/s. A client that was
idle for 2 s has refilled to the cap and can fire a 200-request burst instantly, then settles
back to 100/s. Try to hold 50,000/s — like Adda's runaway bot — and they get one burst of 200,
then 429s for everything above 100/s. The database never sees the flood.
A Redis-backed check adds ~0.2–1 ms per request and one small round trip. That's the tax Adda pays on every call for fleet-wide accuracy; it's cheap next to the outage it prevents. One Redis node handles well over 100,000 limit checks per second.
When to use it
Local is fast but leaky; distributed is accurate but costs a round trip
Per-server in-memory limiting is nanoseconds fast but wrong across a fleet — Adda's 10 gateway servers mean a client gets 10× their limit. A shared Redis counter is accurate but adds latency and a hard dependency (if Redis is down, do you fail open and allow everything, or fail closed and block everyone?). A common compromise: a generous local limit as a cheap first pass, backed by a precise distributed limit.
Choose a fair key and return good errors
Limiting by IP punishes everyone behind one NAT or corporate proxy; limiting by API key or user
ID is fairer. Always return 429 with a Retry-After and, ideally, X-RateLimit-Remaining, so
well-behaved clients back off gracefully instead of retrying blindly and making it worse.
Practice
Recap
- Rate limiting caps how much any one client can consume so a single misbehaving caller can't starve everyone sharing the service.
- Token bucket is the default: refill rate
Rsets the sustained cap, bucket sizeBsets the burst allowance, and it avoids the fixed-window boundary loophole. - Across a fleet, keep the count in a shared atomic store, pick a fair key (API key over IP),
return
429withRetry-After, and decide fail-open vs fail-closed on purpose.
The API Gateway
The single front door where Adda's rate limits are enforced.
Idempotency
Handling the retries that a 429 tells clients to make.
Backpressure & Circuit Breakers
Shedding load when limits alone aren't enough.
In an interview
How to discuss this
Name token bucket and explain its two dials — refill rate and burst size — then immediately
raise the distributed problem: per-server counters multiply the limit by the fleet, so you need
a shared atomic store like Redis. Mention returning 429 with Retry-After, choosing a fair key
over raw IP, and the fail-open vs fail-closed decision when the limiter's store is down. That
progression — algorithm, then distribution, then failure mode — is exactly what they're probing
for.
How is this guide?
Last updated on
Publish / Subscribe
One new Adda post must ripple to the feed, search, notifications, and analytics. If the Posts service has to call each of them, it becomes everyone's bottleneck. Pub/sub lets it broadcast and forget.
Idempotency
Fahim taps Post, the network hiccups, his client retries — and the post appears twice. Worse when it's an Adda Premium payment. Idempotency makes "do this once" survive retries.