Handling Claude API 429s in Next.js: what the SDK retries, and what reaches your users
What the SDK already retries on a 429, what it cannot retry once tokens are flowing, and how to show users a clean message instead of a stack trace.
There are two kinds of 429 in an AI product, and they get confused constantly. One is yours: the limiter you put in front of your own routes, which I covered in rate limiting LLM API routes. This post is about the other one — Anthropic rate-limiting you, because a traffic spike pushed your account past its limits. That one you do not control, so the job is to absorb it where you can and fail cleanly where you cannot.
What the SDK already does for you
Before writing any retry logic, know what the official TypeScript SDK already handles. In the installed version (0.104.x) the client retries automatically on 408, 409, 429, and any 5xx — which includes the 529 "overloaded" status — and does not retry other 4xx errors like a bad request or an auth failure. It honors the server's retry-after-ms and retry-after headers when present; otherwise it backs off exponentially from half a second, doubling each attempt, capped at 8 seconds, with up to 25% jitter so a fleet of instances does not retry in lockstep.
The default is two retries. The kit raises it by one in the client singleton:
client = new Anthropic({ maxRetries: 3 });
That is the whole retry layer. Wrapping your own retry loop around the SDK call is a common mistake: you multiply attempts (your three times the SDK's three) and stack backoffs, turning a brief spike into a long hang and extra load on an API that is already telling you to slow down.
What retries cannot fix: the stream has started
Retries happen on the request. For a streaming chat response, once the first tokens have reached the browser, a failure mid-generation cannot be transparently retried — the user has already seen half an answer, and silently starting over would duplicate or contradict it. So a streaming route needs a way to say "this failed" after a successful HTTP status, which is why the kit's stream sends a typed error frame rather than relying on the response code:
function clientSafeError(err: unknown): string {
if (err instanceof Anthropic.RateLimitError) {
return "The AI provider is rate-limiting us. Please retry in a moment.";
}
if (err instanceof Anthropic.APIError) {
return `Model request failed (${err.status ?? "network"}).`;
}
return "Something went wrong generating the response.";
}
The full error is logged server-side; only that short, safe string goes to the client as an error event, and the UI shows it on the half-finished message. The client-side half of that — reading error frames, not just status codes — is in SSE vs WebSockets for LLM streaming.
Never leak the provider error
Raw API errors carry request ids, model names, organization details, and sometimes fragments of your prompt. Pass them through and a user's screenshot of a failure becomes a small information leak — and "Error 429 from upstream provider" also tells every competitor who your provider is. Map known error classes to messages a user can act on ("retry in a moment"), keep the detail in your logs, and give everything unrecognized one generic message.
Things that bite
- Retries spend your time budget. Three retries with backoff can add several seconds — longer if the server sends a large
retry-after. That comes out of your route's max duration, so size the timeout knowing the retries happen inside it. - A 429 under steady load is a capacity problem, not an error-handling problem. Retries absorb spikes. If you hit provider limits every afternoon, you need a higher tier, cheaper routing for some traffic, or tighter per-user limits — not more retries.
- Do not retry what the SDK will not. A 400 or 401 is not going to succeed on the next attempt. Fix the request or the key.
- Keep your own 429 distinct. When your limiter fires, say so ("you are sending too fast"); when the provider does, say that instead. Users and your logs both need to know which wall they hit.
Shipwright — the Next.js and Claude starter kit this blog documents — ships the tuned client, the client-safe error mapping, and the typed error frame wired through the stream, so provider hiccups degrade into a clear message rather than a blank screen. Try it in the live demo.