529 is the one Claude error your code did not cause: Anthropic is overloaded. You cannot fix it — you can only absorb it gracefully. That means patient retries with backoff, a fallback model on latency-sensitive paths, and above all, no bursts of immediate retries that amplify the incident.
The error
{
"type": "error",
"error": { "type": "overloaded_error",
"message": "Overloaded" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Provider-side saturation (launch days, regional incidents). It affects all customers at once. | Backoff with jitter; check Anthropic's status page instead of redeploying your application. |
| Your own load spike hits capacity that is already under pressure. | Spread batch jobs out; a ten-minute delay usually clears it. |
| Confusion with 429: a rate limit looks similar in the logs, but the cause is completely different. | 429 means you exceeded your limits (the server is healthy); 529 means the server is overloaded (your quota is fine). Only 429 includes a Retry-After hint. |
| No fallback is defined, so a provider problem reaches the end user. | Define a fallback chain — within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you can survive a complete incident. |
Retrying without amplifying the incident
Treat 529 like a 429 without Retry-After: exponential backoff starting at about 2 seconds, with jitter, capped at 30–60 seconds, give up after around five attempts, and queue the work. Jitter is the important part: without it, all clients return at once and prolong exactly the overload they are trying to escape.
Failing over instead of failing
For latency-sensitive paths, define a fallback chain. On an OpenAI-compatible endpoint, that is a single changed string — no second SDK and no second account:
PREFERRED = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]
def complete(messages):
last = None
for model in PREFERRED:
try:
return client.chat.completions.create(
model=model, messages=messages, max_tokens=800)
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
last = e # saturado — probar el siguiente nivel
raise lastOnly then look at your code
If 529s appear only for one request type while other calls succeed at the same time, it is not a general incident: check whether that path sends unusually large prompts or fires in a tight loop. If instead it affects all calls at once and disappears on its own, it was a capacity issue — and the work belongs in retry and fallback logic, not a refactor.
If you’re calling through Kunavo
Kunavo routes Claude through more than one upstream path, and its multi-model catalog turns cross-provider failover into changing the model name on the same key and the same wallet — the pattern above does not need a second account. Any 529s that still reach you are never billed. Capacity and price are separate questions; for the latter, per-model rates are listed in the Claude pricing.
Frequently asked questions
Is a 529 my fault?
No. It is provider-side capacity. Your only responsibilities are not to amplify the incident (backoff, jitter) and to have a fallback available if the incident lasts longer than your latency budget.
529 vs. 429 — what is the difference?
429 means you exceeded your limits; the server is fine. 529 means the server itself is overloaded; your quota is fine. Both are retryable; only 429 comes with a Retry-After hint.
How long does a 529 episode usually last?
It is not predictable or guaranteed — which is why the correct response is capped backoff plus a queue, not a fixed wait hard-coded into the application. If your path has a latency budget, the fallback takes over instead of waiting.
Are calls that end in 529 billed?
Not through Kunavo: a request that ends in an error is not billed. Under a direct contract, it depends on the billing rules of the relevant provider.
Related guides
- “Error in message stream” in ChatGPT — causes and how to fix it
- Claude Prices in 2026: Pro, Max, the Per-Token API, and Which Is Cheaper
- Claude Code pricing in 2026: subscription or API, how much a month costs, and which pays off
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.