How to Build a Rate-Limited REST API with Node.js and Redis

How to Build a Rate-Limited REST API with Node.js and Redis Backend

Every public API meets the same client eventually: a loop, no backoff, and a few thousand requests a minute aimed at one endpoint. Everyone else then queues behind that traffic. Rate limiting is how you stop it happening — not as something bolted onto a gateway at the last minute, but as part of the contract your API makes with the people calling it.

What follows is a token-bucket limiter you can build in an afternoon: Redis for shared state, a short Lua script for atomicity, headers that tell clients where they stand, a 429 response they can act on, and a load test to check the arithmetic actually holds.

Why a token bucket beats counting in windows

Picture a bucket that holds tokens. It refills at a fixed rate — say ten per second — and every request takes one token out. If the bucket is empty, the request is refused. That gives you two properties at once: an average rate you can reason about, and permission to burst up to the bucket's capacity.

The simpler alternative, a fixed-window counter, is easy to write and easy to misuse. If the window resets on the minute and the limit is 100, a client can send 100 requests at 11:59:59 and another 100 at 12:00:00. Nothing in your logs looks wrong, but the upstream service saw 200 requests in two seconds. A sliding window log avoids that but stores a timestamp per request, which gets expensive at scale.

Token buckets sit in the useful middle. Memory per client is constant, the maths is O(1), and bursts are shaped rather than banned.

Storing the bucket in Redis

Each client needs a key. Something like rate:key:abc123 for API-key traffic, or rate:ip:203.0.113.7 as a fallback. Under that key, store two values: the current token count as a float, and the timestamp of the last refill. A Redis hash works well; a short JSON string is fine too.

The part people get wrong is atomicity. If you read the token count in Node, calculate the refill, then write the result back, that is three network round trips. Two requests arriving together can both read four tokens, both decide they are allowed, and both write three back. Your limiter silently permits twice what you configured.

Redis solves this with a Lua script, which executes as a single unit — no other command interleaves with it. Keep the script small and predictable:

  1. Read the stored tokens and the last refill timestamp.
  2. Work out elapsed time in seconds and add elapsed × refill rate tokens, capped at the bucket capacity.
  3. If at least one token is available, subtract one, allow the request, and return the new count.
  4. If not, deny it and return how many seconds until the next token appears.
  5. Write the new state back, and set a TTL on the key.

That TTL matters. Without it, every one-off visitor leaves a key behind forever. Set it to roughly the time a full bucket takes to refill, plus a margin — a few minutes is usually plenty.

Wiring it into Express as middleware

A single piece of middleware, registered before your routes, keeps the logic in one place. Two details are worth getting right.

Identify the client deliberately

Prefer an API key or authenticated user ID over an IP address. If you do fall back to IP, set trust proxy correctly, or every request behind your load balancer will look like it came from the same address and one client will lock out everyone else.

A limiter is only as strong as the identity you key it on. Per-IP limits are the quickest to build and the easiest to sidestep.

Call the script efficiently

Load the script into Redis once at startup and call it by SHA hash rather than shipping the source on every request. Keep one shared connection for the limiter — a new client per request will exhaust sockets long before your rate limit matters. Handle the case where Redis has been restarted and no longer recognises the hash: reload the script and retry once.

Headers clients can actually use

Send a consistent set on every response, not just the rejections:

  • RateLimit-Limit — the bucket capacity, so clients know the burst ceiling.
  • RateLimit-Remaining — whole tokens left, floored rather than rounded up.
  • RateLimit-Reset — seconds until at least one token is available again.
  • Retry-After — only on a 429, giving the same number as a plain delay.

The IETF draft uses the unprefixed names; plenty of older clients only look for the X-RateLimit- variants. Choose one convention, document it, and if you have existing integrations, send both for a release or two before dropping the older set.

A 429 response that explains itself

Reject with 429 and a JSON body that names the problem: an error code, a short human-readable message, and the retry delay in seconds. Do not close the socket, and do not return a 500 — clients treat those differently, and monitoring tools will page you for what is normal behaviour.

Two refinements are worth considering before you go live. First, if some endpoints are far more expensive than others, let them cost more than one token so the limit reflects real work. Second, roll the limiter out in log-only mode for a week: count what would have been refused without refusing anything. That tells you whether your capacity is generous or absurd before real users find out.

Load testing the limiter

Tools such as autocannon or k6 are enough. Run against staging with the same Redis instance shape as production, and watch for these things:

  • A single client blasting requests should see roughly the bucket capacity succeed, then a steady trickle at the refill rate — not a flat wall of rejections.
  • p99 latency should stay close to your unthrottled baseline. If it climbs, the limiter is adding Redis round trips to every request and you may need to batch or cache.
  • Thousands of distinct clients at once should still behave, which is really a test of your key expiry and Redis memory.
  • Kill Redis mid-test and see what happens. Decide in advance whether you fail open, letting traffic through and logging loudly, or fail closed and reject everything. Most public APIs fail open; anything expensive or sensitive usually should not.

Ship it, then watch it

Because the state lives in Redis rather than in process memory, the limiter works unchanged across every Node instance you run — which is the main reason to accept the extra hop. Add a counter for refusals, broken down by route, so you can see whether the limit is doing its job or simply annoying legitimate users. Then tune the numbers from that data rather than from a guess.

If you later put a CDN or API gateway in front of the service, check where the limiting happens. Two limiters with different numbers is worse than one, and the one closest to the client is usually the one that protects everything behind it.

Photo: kieutruongphoto / Pixabay