Back to guides
Troubleshooting·August 30, 2026·6 min read

529 overloaded_error in the Claude API — what it means and how to handle it

529 is the only Claude error your code did not cause: Anthropic itself is overloaded. You cannot fix it — you can absorb it well. That means patient retries, a fallback model, and never amplifying the incident with immediate attempts.

529 is the only Claude error your code did not cause: Anthropic itself is overloaded. You cannot fix it — you can absorb it well. That means patient retries, a fallback model, and never amplifying the incident with immediate attempts.

The error

Response (HTTP 529)
{
  "type": "error",
  "error": { "type": "overloaded_error",
             "message": "Overloaded" }
}

Causes and fixes at a glance

CauseFix
Provider saturation (launch days, regional incidents)Exponential backoff with jitter. Check the provider's status page instead of redeploying.
Your traffic spike dropped during a partial incidentDistribute batch jobs; ten minutes of waiting usually resolves it.
Immediate retry loopTrying again immediately multiplies the load and prolongs the incident for everyone, including you.

Retry like a good citizen

Treat 529 like a 429 without a retry-after header: exponential backoff starting at ~2s, with jitter, a 30–60s cap, giving up after ~5 attempts and queuing the work. The same code branch that handles 429 works for 529.

retry.py
import time, random
from openai import APIStatusError

def com_retry(fn, tentativas=5):
    for i in range(tentativas):
        try:
            return fn()
        except APIStatusError as e:
            if e.status_code not in (429, 500, 529):
                raise
            espera = min(2 ** i + random.random(), 60)
            time.sleep(espera)
    raise RuntimeError("esgotou as tentativas")

Switch models instead of going down

For latency-sensitive paths, define a fallback: within the same family (Sonnet → Haiku), behavior remains similar; across providers (Claude → GPT), you survive an entire incident. On an OpenAI-compatible endpoint, this is a one-string change.

failover.py
PREFERIDOS = ["claude-sonnet-5", "claude-haiku-4-5", "gpt-5-6-terra"]

def completar(mensagens):
    ultimo = None
    for modelo in PREFERIDOS:
        try:
            return client.chat.completions.create(
                model=modelo, messages=mensagens, max_tokens=800)
        except APIStatusError as e:
            if e.status_code not in (429, 500, 529):
                raise
            ultimo = e          # saturado — tenta o próximo
    raise ultimo

Do not confuse 529 with 429 or 402

429 means you exceeded your limits (the server is fine). 529 means the server is overloaded (your quota is fine). 402 means insufficient balance. They look similar in logs and require completely different fixes: only 429 and 529 should be retried.

If you’re calling through Kunavo

At Kunavo, the same multi-model catalog sits behind one key and one wallet, so failover between providers in the example above is a model-name change — it requires neither a second account nor a second registration. Failed requests are not charged. Capacity and price are separate questions; for the latter, per-token rates are in our Claude API pricing guide.

Frequently asked questions

Is the 529 error my fault?

No. It is capacity on the provider's side. Your only responsibilities are not to amplify the problem (backoff with jitter) and to have somewhere to migrate if the incident lasts longer than your latency budget.

What is the difference between 529 and 429?

429 means you exceeded your limits; 529 means the server is overloaded. Both can be retried, but only 429 usually comes with a retry-after hint.

Will I be charged for a request that returned 529?

It should not — the request produced no tokens. At Kunavo, failed requests are not deducted from your balance.

Related guides

More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.