529 is the only Claude error your code did not cause: Anthropic itself is overloaded. You cannot fix it — you can absorb it well. That means patient retries, a fallback model, and never amplifying the incident with immediate attempts.
The error
{
"type": "error",
"error": { "type": "overloaded_error",
"message": "Overloaded" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Provider saturation (launch days, regional incidents) | Exponential backoff with jitter. Check the provider's status page instead of redeploying. |
| Your traffic spike dropped during a partial incident | Distribute batch jobs; ten minutes of waiting usually resolves it. |
| Immediate retry loop | Trying again immediately multiplies the load and prolongs the incident for everyone, including you. |
Retry like a good citizen
Treat 529 like a 429 without a retry-after header: exponential backoff starting at ~2s, with jitter, a 30–60s cap, giving up after ~5 attempts and queuing the work. The same code branch that handles 429 works for 529.
import time, random
from openai import APIStatusError
def com_retry(fn, tentativas=5):
for i in range(tentativas):
try:
return fn()
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
espera = min(2 ** i + random.random(), 60)
time.sleep(espera)
raise RuntimeError("esgotou as tentativas")Switch models instead of going down
For latency-sensitive paths, define a fallback: within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you survive an entire incident. On an OpenAI-compatible endpoint, this is a one-string change.
PREFERIDOS = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]
def completar(mensagens):
ultimo = None
for modelo in PREFERIDOS:
try:
return client.chat.completions.create(
model=modelo, messages=mensagens, max_tokens=800)
except APIStatusError as e:
if e.status_code not in (429, 500, 529):
raise
ultimo = e # saturado — tenta o próximo
raise ultimoDo not confuse 529 with 429 or 402
429 means you exceeded your limits (the server is fine). 529 means the server is overloaded (your quota is fine). 402 means insufficient balance. They look similar in logs and require completely different fixes: only 429 and 529 should be retried.
If you’re calling through Kunavo
At Kunavo, the same multi-model catalog sits behind one key and one wallet, so failover between providers in the example above is a model-name change — it requires neither a second account nor a second registration. Failed requests are not charged. Capacity and price are separate questions; for the latter, per-token rates are in our Claude API pricing guide.
Frequently asked questions
Is the 529 error my fault?
No. It is capacity on the provider's side. Your only responsibilities are not to amplify the problem (backoff with jitter) and to have somewhere to migrate if the incident lasts longer than your latency budget.
What is the difference between 529 and 429?
429 means you exceeded your limits; 529 means the server is overloaded. Both can be retried, but only 429 usually comes with a retry-after hint.
Will I be charged for a request that returned 529?
It should not — the request produced no tokens. At Kunavo, failed requests are not deducted from your balance.
Related guides
- Claude API 429 rate_limit_error — what it means and how to fix it
- 401 authentication_error / invalid x-api-key — what to check, in order
- Claude API 529 overloaded_error — what it is and how to ride it out
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.