429 means “slow down,” not “you are out of money.” The difference matters: one is fixed by waiting, the other by topping up — and confusing them leads to hours of debugging in the wrong place.
The error
{
"type": "error",
"error": { "type": "rate_limit_error",
"message": "Number of requests has exceeded your rate limit" }
}Causes and fixes at a glance
| Cause | Fix |
|---|---|
| Requests per minute above your account limit | Queue requests and limit client-side concurrency instead of sending everything at once. |
| Tokens per minute above the limit | Long prompts consume the token quota long before the request quota. Reduce the context or split the work. |
| Several processes using the same key | The limit applies to the key, not the process. Parallel workers count against the same quota. |
| Retry without backoff | Retrying immediately keeps you permanently above the limit. Exponential backoff with jitter is mandatory. |
Honor retry-after when it is provided
When the response includes the retry-after header, it is not a suggestion: it is the exact time after which the request will be accepted again. Waiting less guarantees another 429.
import time
from openai import APIStatusError
try:
resp = client.chat.completions.create(model=MODELO, messages=msgs)
except APIStatusError as e:
if e.status_code == 429:
espera = float(e.response.headers.get("retry-after", 5))
time.sleep(espera)
resp = client.chat.completions.create(model=MODELO, messages=msgs)
else:
raiseLimit concurrency at the source
The most common cause is not total volume but a burst: twenty requests sent at the same instant exceed a limit that sixty requests spread across a minute would not. A semaphore solves what retrying alone cannot.
import asyncio
LIMITE = asyncio.Semaphore(4) # no máximo 4 chamadas simultâneas
async def chamar(msgs):
async with LIMITE:
return await client.chat.completions.create(
model=MODELO, messages=msgs)Confirm that it is a limit, not a balance issue
A 429 never means you are out of credits — that is 402. If your logs mix the two, separate them by status before investigating: the fix for 429 is timing, while the fix for 402 is a top-up. Retrying a 402 will fail forever.
If you’re calling through Kunavo
At Kunavo, limits apply per key and the balance is a separate prepaid wallet, so the two cases have different statuses: 429 for rate limits and 402 when the balance does not cover the call — never one disguised as the other. Rejected requests are not billed. The per-token rates, which determine how much each call consumes from your balance, are listed in our Claude API pricing guide.
Frequently asked questions
Does 429 mean my credits are gone?
No. Insufficient balance is 402. 429 is about speed: you sent too many requests or tokens in too short a time, and waiting resolves it.
How long should I wait?
If the retry-after header is provided, exactly that long. Without it, use exponential backoff starting at 1–2 seconds with jitter, up to a cap of 30–60 seconds.
Will increasing the limit fix it?
It helps if the volume is genuinely high, but most 429s come from short bursts. Limiting concurrency usually fixes the issue without changing any limit.
Related guides
- 529 overloaded_error in the Claude API — what it means and how to handle it
- 401 authentication_error / invalid x-api-key — what to check, in order
- Claude API 429 rate_limit_error — causes and the fix that holds
More error semantics live in the error reference; getting a key takes a minute via signing up and the authentication guide.