529 is the only Claude error not caused by your code: Anthropic is overloaded. You cannot fix it; you can only handle it well — patient retries with backoff, fallback models for latency-sensitive paths, and absolutely no immediate retry storm that amplifies the outage.
The error
{
"type": "error",
"error": { "type": "overloaded_error",
"message": "Overloaded" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Provider-side overload (new model launch day, regional outage). All customers encounter it simultaneously. | Wait with jittered backoff; check Anthropic’s status page instead of redeploying your application. |
| Your own traffic spike collided with capacity that was already tight. | Spread batch work over time; usually delaying it by ten minutes is enough. |
| Confusing it with 429. They look similar in logs, but the causes are completely different. | 429 means you exceeded your own limit (server healthy); 529 means the server itself is overloaded (your quota is fine). Only 429 includes a Retry-After hint. |
| No fallback was defined, so the provider’s problem reaches end users unchanged. | Define the fallback order first — same family (Sonnet → Haiku) behaves more similarly; cross-provider (Claude → GPT) can survive a provider-wide outage. |
Prevent retries from amplifying the outage
Handle 529 as “a 429 without Retry-After”: exponential backoff starting at about 2 seconds, add jitter, cap at 30–60 seconds, give up after about five attempts, and queue the work. Jitter is what makes it effective: without it, all clients return simultaneously and prolong the congestion they are trying to escape.
Do not go down; route around it
Define the fallback chain in advance for latency-sensitive paths. On an OpenAI-compatible endpoint, this means changing one string — no second SDK or account is needed:
PREFERRED = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]
def complete(messages):
last = None
for model in PREFERRED:
try:
return client.chat.completions.create(
model=model, messages=messages, max_tokens=800)
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
last = e # 過載 —— 試下一個
raise lastOnly then suspect your own code
If only one request type returns 529 while other calls succeed at the same time, it is not a total outage: check whether that path sends unusually large prompts or sends repeatedly in a short loop. Conversely, if all calls start returning 529 simultaneously and recover on their own later, the cause is capacity — change retries and fallbacks, not the architecture.
If you’re calling through Kunavo
Kunavo distributes Claude across more than one upstream path, and its multi-model catalog makes cross-provider fallback a matter of “same key, same balance, change only the model name” — the code above needs no second account. Even if a 529 reaches you, it is never charged. Capacity and price are separate questions; for the latter, model rates are listed in the Claude API pricing table.
Frequently asked questions
Is 529 my problem?
No. It is a provider-side capacity problem. Your responsibilities are only two: do not amplify the outage (backoff and jitter), and have somewhere to route requests when the outage exceeds your latency budget.
What is the difference between 529 and 429?
429 means you exceeded your own limit while the server is healthy; 529 means the server itself is overloaded while your quota is fine. Both can be retried, but only 429 includes a Retry-After hint.
How long does 529 usually last?
It cannot be predicted or guaranteed — so the correct answer is bounded backoff plus a queue, not a hard-coded wait time. If the path has a latency budget, a fallback should take over instead of waiting.
Are calls that fail with 529 charged?
Not through Kunavo: requests ending in an error are not billed. With a direct provider contract, it depends on that provider’s billing rules.
Related guides
- Why ChatGPT Says “Message Stream Error” and How to Fix It
- Claude pricing 2026 — subscription prices, API rates, and the break-even point
- Claude Code Costs 2026 — Subscription and API Pay-as-You-Go, Actual Amounts and Break-Even Point
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.