Back to guides
Troubleshooting·September 15, 2026·6 min read

Claude API error 529 overloaded_error — what it is and how to absorb it

529 is the one Claude error your code did not cause: Anthropic is overloaded. You cannot fix it — you can only absorb it gracefully. That means patient retries with backoff, a fallback model on latency-sensitive paths, and above all, no bursts of immediate retries that amplify the incident.

529 is the one Claude error your code did not cause: Anthropic is overloaded. You cannot fix it — you can only absorb it gracefully. That means patient retries with backoff, a fallback model on latency-sensitive paths, and above all, no bursts of immediate retries that amplify the incident.

The error

Response (HTTP 529)
{
  "type": "error",
  "error": { "type": "overloaded_error",
             "message": "Overloaded" }
}

Causes and fixes at a glance

CauseFix
Provider-side saturation (launch days, regional incidents). It affects all customers at once.Backoff with jitter; check Anthropic's status page instead of redeploying your application.
Your own load spike hits capacity that is already under pressure.Spread batch jobs out; a ten-minute delay usually clears it.
Confusion with 429: a rate limit looks similar in the logs, but the cause is completely different.429 means you exceeded your limits (the server is healthy); 529 means the server is overloaded (your quota is fine). Only 429 includes a Retry-After hint.
No fallback is defined, so a provider problem reaches the end user.Define a fallback chain — within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you can survive a complete incident.

Retrying without amplifying the incident

Treat 529 like a 429 without Retry-After: exponential backoff starting at about 2 seconds, with jitter, capped at 30–60 seconds, give up after around five attempts, and queue the work. Jitter is the important part: without it, all clients return at once and prolong exactly the overload they are trying to escape.

Failing over instead of failing

For latency-sensitive paths, define a fallback chain. On an OpenAI-compatible endpoint, that is a single changed string — no second SDK and no second account:

failover.py
PREFERRED = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]

def complete(messages):
    last = None
    for model in PREFERRED:
        try:
            return client.chat.completions.create(
                model=model, messages=messages, max_tokens=800)
        except APIStatusError as e:
            if e.status_code not in (429, 500, 529):
                raise
            last = e          # saturado — probar el siguiente nivel
    raise last

Only then look at your code

If 529s appear only for one request type while other calls succeed at the same time, it is not a general incident: check whether that path sends unusually large prompts or fires in a tight loop. If instead it affects all calls at once and disappears on its own, it was a capacity issue — and the work belongs in retry and fallback logic, not a refactor.

If you’re calling through Kunavo

Kunavo routes Claude through more than one upstream path, and its multi-model catalog turns cross-provider failover into changing the model name on the same key and the same wallet — the pattern above does not need a second account. Any 529s that still reach you are never billed. Capacity and price are separate questions; for the latter, per-model rates are listed in the Claude pricing.

Frequently asked questions

Is a 529 my fault?

No. It is provider-side capacity. Your only responsibilities are not to amplify the incident (backoff, jitter) and to have a fallback available if the incident lasts longer than your latency budget.

529 vs. 429 — what is the difference?

429 means you exceeded your limits; the server is fine. 529 means the server itself is overloaded; your quota is fine. Both are retryable; only 429 comes with a Retry-After hint.

How long does a 529 episode usually last?

It is not predictable or guaranteed — which is why the correct response is capped backoff plus a queue, not a fixed wait hard-coded into the application. If your path has a latency budget, the fallback takes over instead of waiting.

Are calls that end in 529 billed?

Not through Kunavo: a request that ends in an error is not billed. Under a direct contract, it depends on the billing rules of the relevant provider.

Related guides

More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.