529 is the one Claude error your code did not cause: Anthropic itself is overloaded. You cannot fix it — only handle it cleanly. That means patient retries with backoff, a fallback model for latency-sensitive paths, and absolutely no immediate retry storms that make the incident worse.
The error
{
"type": "error",
"error": { "type": "overloaded_error",
"message": "Overloaded" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Provider-side load (launch days, regional incidents). It affects all customers simultaneously. | Backoff with jitter; check Anthropic's status page instead of redeploying your application. |
| Your own load spike hits capacity that is already under pressure. | Stagger batch jobs; a ten-minute delay usually clears it. |
| Confusion with 429: a rate limit looks similar in logs, but the cause is completely different. | 429 means you exceeded your limits (the server is healthy); 529 means the server is overloaded (your quota is fine). Only 429 provides a Retry-After hint. |
| No fallback is defined, so a provider problem reaches the end user. | Define a fallback chain — within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you can survive a complete outage. |
Retrying without amplifying the incident
Treat 529 like a 429 without Retry-After: exponential backoff starting at about 2 seconds, with jitter, capped at 30–60 seconds, give up after around five attempts, and put the work in a queue. Jitter is the key part: without it, all clients return at the same time and prolong exactly the overload they are trying to escape.
Fail over instead of failing
For latency-sensitive paths, define a fallback chain. On an OpenAI-compatible endpoint, that is one changed string — no second SDK, no second account:
PREFERRED = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]
def complete(messages):
last = None
for model in PREFERRED:
try:
return client.chat.completions.create(
model=model, messages=messages, max_tokens=800)
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
last = e # überlastet — nächste Stufe versuchen
raise lastOnly then inspect your own code
If 529 occurs only for a single request type while other calls go through at the same time, it is not a provider-wide incident: check whether that path sends unusually large prompts or fires in a tight loop. If, on the other hand, it affects all calls simultaneously and then disappears on its own, it was a capacity issue — so the work belongs in retry and fallback, not refactoring.
If you’re calling through Kunavo
Kunavo routes Claude through more than one upstream path, and its multi-model catalog makes cross-provider failover a changed model name on the same key and the same billing account — the pattern above does not need a second account. Any 529s that still reach you are never billed. Capacity and price are two separate questions; for the latter, per-model rates are listed in the Claude API pricing.
Frequently asked questions
Is a 529 my fault?
No. It is provider-side capacity. Your responsibility is limited to not amplifying the incident (backoff, jitter) and having a fallback available if the incident lasts longer than your latency budget.
529 or 429 — what is the difference?
429 means you exceeded your limits; the server is healthy. 529 means the server itself is overloaded; your quota is fine. Both are retryable; only 429 includes a Retry-After hint.
How long does a 529 episode usually last?
It is not predictable and cannot be guaranteed — which is why the correct response is capped backoff plus a queue, not a wait time hard-coded into the application. If your path has a latency budget, the fallback takes over instead of waiting.
Are failed 529 calls billed?
Not through Kunavo: a request that ends in an error is not billed. Under a direct contract, it depends on the billing rules of the relevant provider.
Related guides
- “Error in message stream” in ChatGPT — causes and how to fix it
- Claude Prices 2026 — What Does Claude Cost? Subscription, API, and Claude Code
- Gemini API Prices 2026 — Per-Model Rates, Around 70% Below Google’s List Prices
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.