Node.js & APIs · 18 min · 150 XP
Rate limiting and 429s
Limit how fast one caller can hit an endpoint with a token bucket, and tell them when to come back with a 429.
Any endpoint that does real work can be called faster than you can afford. Some callers are attackers trying passwords; most are a client stuck in a retry loop, a script someone forgot, or a page that re-fetches on every keystroke. Rate limiting caps how many requests one caller gets in a period. It matters most on endpoints that are expensive (search, AI calls, exports) or sensitive (login, password reset, sending email or SMS).
The simplest limiter counts requests per fixed window: 100 per minute, reset on the minute. It has a known flaw: a caller can send 100 at 12:00:59 and 100 more at 12:01:00, which is 200 in two seconds. A token bucket smooths this. Each caller has a bucket holding up to capacity tokens, refilled at a steady rate. Each request spends one token; with none left, it's refused. The capacity allows a short burst, the refill rate sets the long-run speed, and an idle caller can never bank more than one full bucket.
t=0.0s tokens 3 → request ✓ 2 → ✓ 1 → ✓ 0 → ✗ (retry after 1s) t=1.5s tokens 1.5 → request ✓ 0.5 → ✗ (retry after 1s) t=60s tokens 3 (capped; an idle minute doesn't bank 60)
A refused request gets 429 Too Many Requests with a Retry-After header saying how many seconds to wait. Well-behaved clients wait that long instead of hammering you. Decide what one caller is: the user id when they're signed in, the IP address when they're not. Remember that a whole office or mobile network can share one IP, so anonymous limits need to be generous.
A limiter in one process's memory only works with one process. Run three server instances behind a load balancer and each one has its own buckets, so every caller gets three times the limit. Production limiters keep their counters in a shared store such as Redis, or at the edge in your CDN or API gateway.
Loading your workspace…