Errors & rate limits
Standard HTTP status codes; 429 carries Retry-After.
| Status | Meaning |
|---|---|
| 400 | Bad request — malformed body or missing model/messages. |
| 401 | Unauthorized — missing or invalid API key. |
| 404 | Not found — unknown model id. |
| 429 | Rate limited — slow down; honor Retry-After. |
| 500 | Server error — retry with backoff. |
Rate limits
Your account is limited on one dimension: how many requests run at the same time. There is no per-minute request or token cap. The limit is shared across all your keys, and exceeding it returns 429 with Retry-After.
- 20 concurrent requests, shared across all your keys
- No requests-per-minute or tokens-per-minute cap
Determinism
Responses are not bit-for-bit reproducible: GPU batching and floating-point numerics vary between runs, so identical requests can return different completions even with identical parameters. Design retries, caching, and tests accordingly.
API compatibility notes
- The user parameter is accepted but ignored.
- The suffix parameter of the Completions API is not supported.
- The n parameter must be 1. Higher values are rejected with a 400 rather than ignored.
- best_of is not supported and is silently ignored.