Rate limits
Requests per minute, per token
Limits are counted per token, not per organization or per person: two tokens have separate budgets.
Knowing where you stand
Successful responses carry two headers:
Slow down as X-RateLimit-Remaining approaches zero.
When you go over
A request over the limit answers 429:
Wait before retrying: the budget refills within a minute. Retry with an increasing delay rather than immediately, and never retry in a tight loop.
The request quota of the organization is a separate limit, counted over the billing period: see Organizations.
