Rate limiting LLM API routes in Next.js: the limiter, the key, and the Redis seam
A fixed-window limiter keyed by user or IP, wired before the model call to return 429, with a Redis seam for serverless, and what it actually bounds.
A monthly quota caps what a user spends over a billing period. It does nothing about the next ten seconds. Without a rate limit, one client can fire a hundred concurrent requests before the quota check has caught up — running your bill up in a burst, tripping the provider's own rate limits, or just starving the event loop. A rate limit is the second ceiling, and it belongs in front of every route that calls the model.
A fixed-window limiter, keyed by identity
The mechanism is small: a counter per key that resets on a fixed interval. In-memory, that is a map of buckets:
export function rateLimit(key: string, opts?: { limit?: number; windowMs?: number }) {
const limit = opts?.limit ?? 20;
const windowMs = opts?.windowMs ?? 60_000;
const now = Date.now();
let bucket = buckets.get(key);
if (!bucket || bucket.resetAt <= now) {
bucket = { count: 0, resetAt: now + windowMs };
buckets.set(key, bucket);
}
bucket.count += 1;
return { ok: bucket.count <= limit, remaining: Math.max(0, limit - bucket.count), resetAt: bucket.resetAt };
}
What matters more than the algorithm is the key. Limit per identity, not globally: a single global counter lets one abuser exhaust the window for everyone. Use the authenticated user id when the caller is signed in, and fall back to the IP for anonymous traffic — the same identity split the quota and demo attribution use, so signed-in users get their own budget and the public demo is still bounded per address.
Wire it before the model call
The limiter runs first in the route, ahead of body validation and the quota check — reject the cheap way before doing any work, and return 429 with a Retry-After so a well-behaved client backs off instead of hammering:
const limit = rateLimit(`chat:${userId}`);
if (!limit.ok) {
return Response.json(
{ error: "Rate limit exceeded. Try again in a minute." },
{ status: 429, headers: { "Retry-After": String(Math.ceil((limit.resetAt - Date.now()) / 1000)) } },
);
}
The Redis seam
In-memory is exactly right for a single Node process — local dev, or one long-lived container. It quietly breaks on serverless or any multi-instance deploy, because each instance keeps its own map: with four instances behind a load balancer, your "20 per minute" is really up to 80. The fix is not to rewrite callers; it is to keep this same function signature and swap the internals for a shared store like Upstash Redis. Designing the seam in from the start — one function every route calls — means that change is one file, not a migration.
What a rate limit actually bounds
Concretely: worst-case spend per identity per window is limit × max-cost-per-request. That is the number a rate limit gives you that a quota cannot — a ceiling on the burst, where the quota ceilings the month. Together they bound both the second and the billing period; ship only one and the other door is open. This is the pairing behind keeping a free tier from bankrupting you.
Things that bite
- Fixed windows allow an edge burst. A client can send a full window at the end of one interval and another at the start of the next — briefly 2x. Fine for spend control; use a sliding window if you need a hard guarantee.
- In-memory counters reset on cold start. A freshly spun-up serverless instance has empty buckets, so limits are softer than they look right after a deploy — another reason the shared store matters in production.
- IP keys collide. Corporate NATs and mobile carriers put many users behind one address; prefer the user id whenever the request is authenticated so you do not throttle a whole office as one.
Shipwright — the Next.js and Claude starter kit this blog documents — ships this limiter behind a one-line call in the chat route, keyed by identity and ready to swap for Redis, alongside the per-request metering and quota it works with. See it holding the line in the live demo.