Rate limits
Limits keep the platform responsive for everyone and protect you from runaway spend. They are set per plan and apply to the account, across all of its API keys.
Limits by plan#
| Plan | Requests / min | Tokens / min | Concurrent requests | Max context | API keys |
|---|---|---|---|---|---|
| Pay as you go | 60 | 200,000 | 4 | 128K | 5 |
| Starter | 60 | 300,000 | 4 | 128K | 5 |
| Pro | 300 | 1,000,000 | 10 | 256K | 20 |
| Power | 1,000 | 4,000,000 | 30 | 256K | 50 |
| Business | 3,000 | 15,000,000 | 100 | 256K | 200 |
Current values for your account are returned by GET /v1/account and shown on the Usage page. Individual models may carry lower limits during capacity incidents; the effective limit is always the lower of the two.
Tokens per minute counts the estimated prompt size plus the expected output (up to 4,096 tokens) at the time the request starts, so a burst of long prompts can hit the token limit well before the request limit.
Max context caps the prompt size regardless of the model's own window. A pay-as-you-go or Starter request to a 256K model is still limited to 128K tokens and fails with context_length_exceeded above that.
Reading the headers#
Every chat completion response includes the current state of your request and token windows:
x-ratelimit-limit-requests: 300
x-ratelimit-remaining-requests: 287
x-ratelimit-limit-tokens: 1000000
x-ratelimit-remaining-tokens: 942115
x-ratelimit-reset-requests: 41s
x-request-id: req_01J9X3Q5K7M2N8P4R6T0V2W4Y6When a limit is exceeded the response is 429 with a retry-after header (in seconds) and one of these codes:
code | Trigger | Typical retry-after |
|---|---|---|
rate_limit_exceeded | Requests per minute | Seconds until the window resets |
tokens_per_minute_exceeded | Tokens per minute | Seconds until the window resets |
concurrency_limit_exceeded | In-flight requests | 2 |
daily_limit_exceeded | Plan's daily token cap (if any) | Seconds until UTC midnight |
monthly_limit_exceeded | Plan's monthly token cap (if any) | Seconds until UTC midnight |
{
"error": {
"message": "Rate limit reached: 300 requests per minute. Retry in 41s.",
"type": "rate_limit_error",
"code": "rate_limit_exceeded",
"request_id": "req_01J9X3Q5K7M2N8P4R6T0V2W4Y6"
}
}Rate-limited requests are never billed.
Other limits#
| Limit | Value |
|---|---|
| Request body size | 2 MB |
| Messages per request | 200 |
| Text content per message | 1,000,000 characters |
| Tools per request | 128 |
| API key creation | 20 per hour |
Staying under the limits#
The official OpenAI SDKs retry 429 and 5xx responses with exponential back-off and honour retry-after automatically. Set maxRetries (JavaScript) or max_retries (Python) to tune it.
- Respect
retry-after. Retrying sooner just consumes more of the window. - Use a client-side semaphore equal to your plan's concurrency limit. Excess parallelism turns into
concurrency_limit_exceededand wasted round trips. - Stream long responses. Streaming does not change the limits but lets you show progress instead of waiting.
- Keep prompts tight. Fewer prompt tokens means more requests per token window and a lower bill.
- Watch
x-ratelimit-remaining-*and pre-emptively slow down when they approach zero rather than waiting for a429.
If you consistently need more headroom, upgrade your plan from Billing. Limits apply immediately after the plan changes.