How AI API Rate Limits Work

An AI API rate limit is a provider-enforced restriction on how quickly you can send requests or consume resources such as input and output tokens. Limits help providers protect shared infrastructure, allocate capacity fairly, and prevent accidental loops or traffic spikes from overwhelming a service.
How AI API Rate Limits Work

What is an AI API rate limit?

An AI API rate limit controls the amount of work an account, project, organization, model, endpoint, or application can submit during a period of time. Depending on the service, the limit may count requests, tokens, concurrent jobs, images, audio duration, queued batch tokens, spending, or total usage over a daily or rolling window.

For example, a provider might limit both the number of requests per minute and the number of tokens processed per minute. These are different constraints. Many short requests may exhaust the request limit first, while a few requests containing long documents may exhaust the token limit.

Rate limits are not the same as a model's output format, context window, or maximum response length. They control usage over time rather than what can fit into one request.

Why providers impose rate limits

AI inference requires shared computing resources, and requests can vary greatly in cost. A short classification request and a long document analysis may each count as one request, but the second can use far more tokens and processing capacity.

Rate limits help providers:

  • Protect infrastructure from traffic spikes, accidental loops, and abuse.
  • Share limited inference capacity among customers.
  • Control usage at different account, project, model, and endpoint levels.
  • Offer predictable service tiers without promising unlimited capacity.

They also protect your own application. Without an application-level limit or queue, one user, scheduled job, or software bug could consume an entire shared allowance.

The main dimensions of rate limiting

Requests per minute

Requests per minute, or RPM, limits how many API calls can be made during a minute or comparable rolling period. Requests per day, commonly called RPD, apply the same idea over a longer period.

A displayed limit of 60 requests per minute does not necessarily mean that 60 requests can be sent simultaneously. Providers may also enforce shorter sub-windows, burst limits, concurrency limits, or ramp-rate rules.

Tokens per minute

Tokens per minute, or TPM, limits the amount of model text processed over time. A token is a unit produced by a model's tokenizer; it is not exactly the same as a word. Token counts vary with language, formatting, and code.

Token accounting differs between providers. A service may count input tokens, output tokens, estimated output capacity, or other provider-specific usage. Some systems may reserve capacity based on the requested maximum output before generation finishes. Check the provider's documentation rather than assuming that all APIs count tokens identically.

Concurrency

Concurrency is the number of requests or jobs allowed to be active at once. It is different from RPM. An application may remain below its per-minute request limit while still opening more simultaneous generations than the provider allows.

Quotas, spend limits, and daily allowances

A quota is a broader usage allowance. It may cover daily or monthly usage, queued batch tokens, or an account's approved capacity. A spend limit controls money spent rather than request speed. Both can prevent requests, but waiting briefly may not resolve an exhausted billing or usage allowance.

Bursts and ramp rates

A burst is a sudden concentration of traffic. A ramp-rate restriction controls how quickly usage increases. An application can have a compliant average rate and still be slowed or rejected after an abrupt traffic jump.

What happens when a request is sent?

  1. The client sends an authenticated request to an endpoint.
  2. The provider identifies relevant scopes such as the account, organization, project, API key, model, region, endpoint, and sometimes user or IP address.
  3. The provider checks one or more counters or budgets, including request rate, token throughput, concurrency, daily quota, and spend.
  4. If capacity is available, the request is admitted and usage accounting is updated.
  5. If a limit is exceeded, the provider may reject the request, queue it, slow it down, or return an overload response.

Providers may use token buckets, leaky buckets, fixed or sliding windows, rolling windows, admission control, concurrency limits, or combinations of these techniques. Public documentation often describes the observable limit without revealing the exact internal algorithm.

Why one limit can be exhausted before another

Imagine an application with a limit of 60 requests per minute and 100,000 tokens per minute. Sixty tiny requests might exhaust the request allowance while using relatively few tokens. Ten large document requests could exhaust the token allowance before the request counter reaches 60.

If the same service also permits only five active requests, sending six at once may cause a concurrency failure or queue even when both per-minute counters have room. The first exhausted dimension determines what happens.

What does HTTP 429 mean?

HTTP 429, “Too Many Requests,” is the standard status commonly used when a client exceeds a rate or resource limit. It is an important signal, but it is not a complete diagnosis.

A 429 might indicate temporary throttling, too many requests, too many tokens, an exhausted usage tier, a project quota, or a spending condition. Some providers also use 429 for non-transient account restrictions. Read the response body, provider-specific error code, response headers, dashboard information, and reset guidance.

A temporary customer throttle is also different from provider overload. Overload may be reported with HTTP 503, HTTP 529, or another provider-specific code even when your own quota is available.

Retry-After and response headers

A response may include a Retry-After header indicating how long to wait before trying again. Not every provider or error includes it. Some APIs also expose remaining request or token capacity and reset times in provider-specific headers.

These headers are useful signals, but their names and meanings are not universal. Treat them according to the provider's documentation.

How to handle rate limits reliably

Use a queue and control concurrency

Instead of allowing every worker to call the API immediately, place work in a queue and process it with a controlled number of workers. This limits simultaneous requests and makes backlog visible.

Pace requests proactively

A rate limiter can spread requests across time rather than sending them in bursts. For token-heavy workloads, request-based pacing alone is insufficient; estimate token demand as well.

Retry transient failures carefully

For a transient 429 or suitable 5xx response, wait before retrying. If Retry-After is present, honor it. Otherwise, use exponential backoff: increase the delay after each failed attempt, add random jitter, and set a maximum delay and retry count.

Jitter prevents many workers from retrying at exactly the same moment. Immediate synchronized retries can create a thundering-herd effect and prolong throttling.

Do not retry indefinitely. Authentication failures, invalid parameters, unsupported models, exhausted credits, hard spend limits, and exhausted daily quotas generally require configuration, billing, or operator action rather than more attempts.

Make repeated work safe

Retries can duplicate work or side effects. For operations that create records, send messages, or trigger external actions, use idempotency keys, deduplication, or durable job state when supported by the application and provider.

Reduce token pressure

Trim unnecessary context, limit output to realistic needs, use concise prompts, and avoid repeatedly sending information that the application does not need. Caching repeated context where supported can also reduce work, but it does not remove every request or concurrency limit.

Use batching for non-urgent work

Batch APIs can be useful for large asynchronous jobs such as classifying records overnight. They often have separate quotas and queue rules, so batching is not a universal escape from provider limits. It trades immediate responses for throughput and may introduce file, queue, processing-time, and model restrictions.

Provider limits can apply at several levels

A provider may enforce limits at the API key, project, organization, account, model, model family, endpoint, region, or application level. Several scopes can apply to one request.

For example, OpenAI documents organization and project limits that vary by model, with some model families sharing limits. Google Gemini documents project-level, model-dependent dimensions such as requests per minute, input tokens per minute, requests per day, and other usage or spend limits. Anthropic documents organization-level usage tiers and uses HTTP 429 for rate-limit conditions, while separately distinguishing overload responses.

These examples show why provider behavior should not be treated as universal. Adding API keys may not increase capacity if the real limit applies to the shared project, organization, account, or billing scope.

  • Context window: limits the input and output that fit in one request. It does not determine how many requests can be sent over time.
  • Maximum output tokens: limits one response. TPM limits aggregate token throughput across requests.
  • Daily or monthly quota: limits total usage over a longer period. Waiting a few seconds may not restore it.
  • Spend limit: controls money spent, not request speed.
  • Provider overload: means the backend may not have enough capacity, even if your quota is available.
  • Application limit: a limit your own software imposes per user, tenant, or endpoint. It can be stricter than the provider's limit.
  • Consumer message cap: a restriction in a chatbot product that should not be assumed to match API behavior.

A practical debugging checklist

When requests begin failing, investigate systematically:

  1. Record the HTTP status, provider error code, error message, timestamp, model, endpoint, and relevant response headers.
  2. Check whether the failure is request-rate, token, concurrency, daily, spend, billing, authorization, or provider-overload related.
  3. Compare estimated input and output tokens with actual usage where available.
  4. Look for burst traffic, synchronized retries, scheduled jobs, worker increases, or a recent ramp in volume.
  5. Check the provider dashboard and current model-specific documentation.
  6. Stop retrying errors that require billing, access, quota, or configuration changes.
  7. Test again with one request at a time, then increase concurrency gradually.

A useful experiment is to send 20 small requests and then five large requests while logging token estimates, status codes, retry delays, concurrency, and elapsed time. Compare evenly paced traffic with a burst. This can reveal whether the workload is request-limited, token-limited, concurrency-limited, or affected by burst behavior.

Common misconceptions

“Below RPM means the request cannot be limited.”

Token, concurrency, daily, project, spend, shared-model, and capacity limits can still apply.

“A 429 always means retry immediately.”

Some 429 responses represent billing, usage-tier, or quota conditions that will not be fixed by retrying. Inspect the error details first.

“Streaming avoids rate limits.”

Streaming changes how output is delivered. It generally does not remove request, token, concurrency, or account-level limits.

Paid tiers may increase allowances or provide access to higher-capacity arrangements, but they still have documented or negotiated constraints.

“Separate API keys create separate quotas.”

Many providers apply limits above the key level, so several keys may draw from the same project, organization, account, or billing pool.

Key takeaways

AI API rate limits are multidimensional controls, not simply a counter of calls. A request can fail because of request volume, token throughput, concurrency, burst behavior, daily quota, spending, or provider capacity.

Reliable applications combine proactive pacing with reactive error handling. Use queues, concurrency limits, token-aware budgeting, exponential backoff with jitter, bounded retries, and clear handling for permanent billing or authorization errors. Finally, treat exact limits and error behavior as provider- and model-specific details that should be checked before production use.


Answers to Frequently Asked Questions

Why can an AI API request fail even when the requests-per-minute limit has not been reached?
AI APIs often enforce multiple limits at once. A request can fail because of token throughput, concurrency, burst behavior, daily or monthly quotas, spend limits, shared project or organization capacity, or provider overload even when the requests-per-minute allowance remains available.
How should applications handle AI API rate limits?
Use a queue, control concurrency, pace requests proactively, estimate token demand, and retry suitable transient errors with exponential backoff and random jitter. Honor the Retry-After header when available, set maximum retry limits, and avoid retrying permanent errors such as authentication failures, invalid parameters, exhausted credits, or hard quotas.
What does HTTP 429 mean when using an AI API?
HTTP 429, “Too Many Requests,” usually indicates that a rate or resource limit was exceeded. It may relate to request volume, token usage, concurrency, quotas, usage tiers, or spending conditions. Check the response body, error code, headers, dashboard, and provider documentation before retrying.
What is an AI API rate limit?
An AI API rate limit controls how much work an account, project, model, endpoint, or application can submit during a defined period. Limits may count requests, tokens, concurrent jobs, images, audio duration, queued batch tokens, spending, or total usage.
What are the main types of AI API rate limits?
Common dimensions include requests per minute (RPM), tokens per minute (TPM), concurrency, daily or monthly quotas, spend limits, burst limits, and ramp-rate restrictions. Several limits can apply to the same request, and the first exhausted limit determines what happens.