How AI API Rate Limits Work
What is an AI API rate limit?
An AI API rate limit controls the amount of work an account, project, organization, model, endpoint, or application can submit during a period of time. Depending on the service, the limit may count requests, tokens, concurrent jobs, images, audio duration, queued batch tokens, spending, or total usage over a daily or rolling window.
For example, a provider might limit both the number of requests per minute and the number of tokens processed per minute. These are different constraints. Many short requests may exhaust the request limit first, while a few requests containing long documents may exhaust the token limit.
Rate limits are not the same as a model's output format, context window, or maximum response length. They control usage over time rather than what can fit into one request.
Why providers impose rate limits
AI inference requires shared computing resources, and requests can vary greatly in cost. A short classification request and a long document analysis may each count as one request, but the second can use far more tokens and processing capacity.
Rate limits help providers:
- Protect infrastructure from traffic spikes, accidental loops, and abuse.
- Share limited inference capacity among customers.
- Control usage at different account, project, model, and endpoint levels.
- Offer predictable service tiers without promising unlimited capacity.
They also protect your own application. Without an application-level limit or queue, one user, scheduled job, or software bug could consume an entire shared allowance.
The main dimensions of rate limiting
Requests per minute
Requests per minute, or RPM, limits how many API calls can be made during a minute or comparable rolling period. Requests per day, commonly called RPD, apply the same idea over a longer period.
A displayed limit of 60 requests per minute does not necessarily mean that 60 requests can be sent simultaneously. Providers may also enforce shorter sub-windows, burst limits, concurrency limits, or ramp-rate rules.
Tokens per minute
Tokens per minute, or TPM, limits the amount of model text processed over time. A token is a unit produced by a model's tokenizer; it is not exactly the same as a word. Token counts vary with language, formatting, and code.
Token accounting differs between providers. A service may count input tokens, output tokens, estimated output capacity, or other provider-specific usage. Some systems may reserve capacity based on the requested maximum output before generation finishes. Check the provider's documentation rather than assuming that all APIs count tokens identically.
Concurrency
Concurrency is the number of requests or jobs allowed to be active at once. It is different from RPM. An application may remain below its per-minute request limit while still opening more simultaneous generations than the provider allows.
Quotas, spend limits, and daily allowances
A quota is a broader usage allowance. It may cover daily or monthly usage, queued batch tokens, or an account's approved capacity. A spend limit controls money spent rather than request speed. Both can prevent requests, but waiting briefly may not resolve an exhausted billing or usage allowance.
Bursts and ramp rates
A burst is a sudden concentration of traffic. A ramp-rate restriction controls how quickly usage increases. An application can have a compliant average rate and still be slowed or rejected after an abrupt traffic jump.
What happens when a request is sent?
- The client sends an authenticated request to an endpoint.
- The provider identifies relevant scopes such as the account, organization, project, API key, model, region, endpoint, and sometimes user or IP address.
- The provider checks one or more counters or budgets, including request rate, token throughput, concurrency, daily quota, and spend.
- If capacity is available, the request is admitted and usage accounting is updated.
- If a limit is exceeded, the provider may reject the request, queue it, slow it down, or return an overload response.
Providers may use token buckets, leaky buckets, fixed or sliding windows, rolling windows, admission control, concurrency limits, or combinations of these techniques. Public documentation often describes the observable limit without revealing the exact internal algorithm.
Why one limit can be exhausted before another
Imagine an application with a limit of 60 requests per minute and 100,000 tokens per minute. Sixty tiny requests might exhaust the request allowance while using relatively few tokens. Ten large document requests could exhaust the token allowance before the request counter reaches 60.
If the same service also permits only five active requests, sending six at once may cause a concurrency failure or queue even when both per-minute counters have room. The first exhausted dimension determines what happens.
What does HTTP 429 mean?
HTTP 429, “Too Many Requests,” is the standard status commonly used when a client exceeds a rate or resource limit. It is an important signal, but it is not a complete diagnosis.
A 429 might indicate temporary throttling, too many requests, too many tokens, an exhausted usage tier, a project quota, or a spending condition. Some providers also use 429 for non-transient account restrictions. Read the response body, provider-specific error code, response headers, dashboard information, and reset guidance.
A temporary customer throttle is also different from provider overload. Overload may be reported with HTTP 503, HTTP 529, or another provider-specific code even when your own quota is available.
Retry-After and response headers
A response may include a Retry-After header indicating how long to wait before trying again. Not every provider or error includes it. Some APIs also expose remaining request or token capacity and reset times in provider-specific headers.
These headers are useful signals, but their names and meanings are not universal. Treat them according to the provider's documentation.
How to handle rate limits reliably
Use a queue and control concurrency
Instead of allowing every worker to call the API immediately, place work in a queue and process it with a controlled number of workers. This limits simultaneous requests and makes backlog visible.
Pace requests proactively
A rate limiter can spread requests across time rather than sending them in bursts. For token-heavy workloads, request-based pacing alone is insufficient; estimate token demand as well.
Retry transient failures carefully
For a transient 429 or suitable 5xx response, wait before retrying. If Retry-After is present, honor it. Otherwise, use exponential backoff: increase the delay after each failed attempt, add random jitter, and set a maximum delay and retry count.
Jitter prevents many workers from retrying at exactly the same moment. Immediate synchronized retries can create a thundering-herd effect and prolong throttling.
Do not retry indefinitely. Authentication failures, invalid parameters, unsupported models, exhausted credits, hard spend limits, and exhausted daily quotas generally require configuration, billing, or operator action rather than more attempts.
Make repeated work safe
Retries can duplicate work or side effects. For operations that create records, send messages, or trigger external actions, use idempotency keys, deduplication, or durable job state when supported by the application and provider.
Reduce token pressure
Trim unnecessary context, limit output to realistic needs, use concise prompts, and avoid repeatedly sending information that the application does not need. Caching repeated context where supported can also reduce work, but it does not remove every request or concurrency limit.
Use batching for non-urgent work
Batch APIs can be useful for large asynchronous jobs such as classifying records overnight. They often have separate quotas and queue rules, so batching is not a universal escape from provider limits. It trades immediate responses for throughput and may introduce file, queue, processing-time, and model restrictions.
Provider limits can apply at several levels
A provider may enforce limits at the API key, project, organization, account, model, model family, endpoint, region, or application level. Several scopes can apply to one request.
For example, OpenAI documents organization and project limits that vary by model, with some model families sharing limits. Google Gemini documents project-level, model-dependent dimensions such as requests per minute, input tokens per minute, requests per day, and other usage or spend limits. Anthropic documents organization-level usage tiers and uses HTTP 429 for rate-limit conditions, while separately distinguishing overload responses.
These examples show why provider behavior should not be treated as universal. Adding API keys may not increase capacity if the real limit applies to the shared project, organization, account, or billing scope.
Rate limits versus related concepts
- Context window: limits the input and output that fit in one request. It does not determine how many requests can be sent over time.
- Maximum output tokens: limits one response. TPM limits aggregate token throughput across requests.
- Daily or monthly quota: limits total usage over a longer period. Waiting a few seconds may not restore it.
- Spend limit: controls money spent, not request speed.
- Provider overload: means the backend may not have enough capacity, even if your quota is available.
- Application limit: a limit your own software imposes per user, tenant, or endpoint. It can be stricter than the provider's limit.
- Consumer message cap: a restriction in a chatbot product that should not be assumed to match API behavior.
A practical debugging checklist
When requests begin failing, investigate systematically:
- Record the HTTP status, provider error code, error message, timestamp, model, endpoint, and relevant response headers.
- Check whether the failure is request-rate, token, concurrency, daily, spend, billing, authorization, or provider-overload related.
- Compare estimated input and output tokens with actual usage where available.
- Look for burst traffic, synchronized retries, scheduled jobs, worker increases, or a recent ramp in volume.
- Check the provider dashboard and current model-specific documentation.
- Stop retrying errors that require billing, access, quota, or configuration changes.
- Test again with one request at a time, then increase concurrency gradually.
A useful experiment is to send 20 small requests and then five large requests while logging token estimates, status codes, retry delays, concurrency, and elapsed time. Compare evenly paced traffic with a burst. This can reveal whether the workload is request-limited, token-limited, concurrency-limited, or affected by burst behavior.
Common misconceptions
“Below RPM means the request cannot be limited.”
Token, concurrency, daily, project, spend, shared-model, and capacity limits can still apply.
“A 429 always means retry immediately.”
Some 429 responses represent billing, usage-tier, or quota conditions that will not be fixed by retrying. Inspect the error details first.
“Streaming avoids rate limits.”
Streaming changes how output is delivered. It generally does not remove request, token, concurrency, or account-level limits.
“Paid access means unlimited capacity.”
Paid tiers may increase allowances or provide access to higher-capacity arrangements, but they still have documented or negotiated constraints.
“Separate API keys create separate quotas.”
Many providers apply limits above the key level, so several keys may draw from the same project, organization, account, or billing pool.
Key takeaways
AI API rate limits are multidimensional controls, not simply a counter of calls. A request can fail because of request volume, token throughput, concurrency, burst behavior, daily quota, spending, or provider capacity.
Reliable applications combine proactive pacing with reactive error handling. Use queues, concurrency limits, token-aware budgeting, exponential backoff with jitter, bounded retries, and clear handling for permanent billing or authorization errors. Finally, treat exact limits and error behavior as provider- and model-specific details that should be checked before production use.
