AI API Errors: Handling 401, 402, 429 and 503
The status code decides the next action. 401 wants a key fix, 402 money, 429 a slower pace, 503 an availability check with an unknown result. Parse status plus error code — never message text.
Short answer
401 means fix the key, 402 money, 429 rate, 503 availability with unknown result. Retry transient only, with backoff and a deadline.
Permanent vs transient
Permanent: 400 (bad request), 401 (key), 402 (balance), 403 (model closed to the key). Retrying them never heals — fix the cause. Transient: 429 and 5xx — retry with exponential backoff, jitter and a global deadline.
Distinguish the two 429s: rate_limit_exceeded slows down, budget_exceeded goes to the dashboard. Blind retries on the second burn money and change nothing.
- 400/401/402/403 — fix the cause
- 429/5xx — backoff with deadline
- two different 429s, different cures
- parse status plus code
Idempotency and diagnostics
Never auto-retry after a partial streamed answer without duplication protection: the server may have finished the job. Put an idempotency key on every side effect and check it before acting.
Minimum investigation set: time, endpoint, model ID, status, error.code, request ID. With this set support finds the operation without your key and prompts.
- no blind retry after partial streams
- idempotency key before acting
- time, endpoint, model, status, code, request ID
Limits that prevent incidents
A per-key spend limit turns a catastrophe into an incident: exhausted means failed requests, not an empty balance. Separate keys per environment show who spends what without digging through a shared bill.
Cap output length per call and set client timeouts: a hanging request is worse than a fast error. After recovery check both balance and key limit — otherwise «money exists, requests rejected».
- spend limit per key before traffic
- separate keys per environment
- cap output, set timeouts
- check balance AND key limit
Sources and related pages
Next step
Russian version: все статьи на русском.