529 is the one Claude error your code did not cause: Anthropic is overloaded. You cannot fix it — only absorb it cleanly. That means patient retries with backoff, a fallback model on latency-sensitive paths, and above all, no bursts of immediate retries that amplify the incident.
The error
{
"type": "error",
"error": { "type": "overloaded_error",
"message": "Overloaded" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Provider-side saturation (launch days, regional incidents). It affects all customers at the same time. | Backoff with jitter; check Anthropic's status page instead of redeploying your application. |
| Your own load spike hits capacity that is already under pressure. | Spread batch processing out; a ten-minute delay is usually enough. |
| Confusion with 429: in the logs, a rate limit looks similar, but the cause has nothing to do with it. | 429 means you exceeded your limits (the server is fine); 529 means the server is overloaded (your quota is fine). Only 429 comes with a Retry-After hint. |
| No fallback is defined, so a provider problem reaches the end user. | Define a fallback chain — within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you can survive a complete incident. |
Retry without amplifying the incident
Treat 529 like a 429 without Retry-After: exponential backoff starting at ~2 seconds, with jitter, capped at 30–60 seconds, give up after approximately five attempts, and queue the work. Jitter is the part that matters: without it, all clients return at the same time and prolong exactly the overload they are trying to escape.
Fail over instead of failing
For latency-sensitive paths, define a fallback chain. On an OpenAI-compatible endpoint, that is a single string to change — no second SDK, no second account:
PREFERRED = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]
def complete(messages):
last = None
for model in PREFERRED:
try:
return client.chat.completions.create(
model=model, messages=messages, max_tokens=800)
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
last = e # saturé — passer au palier suivant
raise lastOnly then inspect your code
If 529s appear only for one request type while other calls succeed at the same time, it is not a general incident: check whether that path sends unusually large prompts or fires in a tight loop. If, instead, it affects all calls at once and then disappears on its own, it was a capacity issue — and the work belongs in retry and fallback, not a rewrite.
If you’re calling through Kunavo
Kunavo routes Claude through more than one upstream path, and its multi-model catalog turns cross-provider failover into simply changing the model name on the same key and the same wallet — the pattern above does not need a second account. Any 529s that still reach you are never billed. Capacity and price are two separate questions; for the latter, per-model rates are listed in the Claude API pricing table.
Frequently asked questions
Is a 529 my fault?
No. It is provider-side capacity. Your only responsibilities are not to amplify the incident (backoff, jitter) and to have a fallback available if the incident lasts longer than your latency budget.
529 or 429 — what is the difference?
429 means you exceeded your limits; the server is fine. 529 means the server itself is overloaded; your quota is fine. Both are retryable; only 429 includes a Retry-After hint.
How long does a 529 period last?
It is not predictable and cannot be guaranteed — which is why the correct response is capped backoff plus a queue, not a delay hard-coded into the application. If your path has a latency budget, the fallback takes over instead of waiting.
Are 529 calls billed?
Through Kunavo, no: a request that ends in an error is not billed. Under a direct contract, it depends on the billing rules of the relevant provider.
Related guides
- “Error in message stream” on ChatGPT — causes and solutions
- Claude API Pricing 2026 — Model Rates, Payment, and Real Costs
- Gemini API Pricing 2026 — Model Rates, Examples, and Cheaper Access
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.