/v1 endpoints share two independent per-minute rate limits:
Both are fixed one-minute windows. Whichever limit you hit first applies. The
defaults are well above normal create-and-poll traffic and may be tuned over
time — read the response headers rather than hard-coding these numbers.
POST /v1/uploads has an additional,
tighter cap of 60 uploads / minute per account — a rolling window — because
it’s free and unmetered. It applies on top of the per-key and per-IP limits
above and returns the same 429 + Retry-After. Regular generation and polling
traffic isn’t affected.Generation jobs also have a separate per-account concurrency cap — the
number of jobs you can have in flight at once, shared with the web workshop.
That cap is independent of the request limits on this page: polling status
never counts against concurrency, and concurrency never shows up in the
X-RateLimit-* headers.The cap is per account (not per key — all your keys and the web app share
it), starts at your plan’s limit (up to 10 on Ultra), and rises
automatically with sustained usage — heavy accounts are granted higher
floors without asking, and we email you when it happens. Your current value is shown under
Billing in the developer console,
where you can also request an increase: tell us your use case and target
concurrency, and the request is reviewed manually — you’ll get the decision by
email, and an approved limit takes effect immediately.Response headers
Every API response reports your current standing:When you exceed a limit
Requests over the limit receive HTTP429 with a Retry-After header (seconds)
and this body:
Other responses that also return 429
Three further conditions answer with429, and none of them is the per-minute
limiter above. Read the body to tell them apart — only the rate limiter and the
Grok cooldown send Retry-After.
Queue full
You have as many generations in flight as your plan’s concurrency cap allows. Waiting a fixed number of seconds is the wrong response: poll your running jobs and submit the next one as they finish. See the concurrency note above for the cap and how to raise it.Monthly spend limit
Your account’s monthly API spend limit has been reached. It is a limit you set yourself, so retrying will not clear it — raise or remove it under Billing in the developer console.Repeated Grok failures
Grok’s content filter rejects some prompts and images, and it is not always consistent about which. After three consecutive failed Grok generations within an hour, further Grok requests pause for ten minutes, growing by ten minutes with each additional failure; any success clears it. Other models are unaffected. Retrying the same prompt will not succeed — change the prompt or the input image, then retry afterretry_after_seconds.
Staying under the limits
- Poll gently. Poll job status every few seconds (images) to every 10–30 seconds (video) — not in a tight loop. See Async jobs.
- Respect
Retry-After. On429, sleep for the advertised seconds, then resume. Add jitter if multiple workers share a key or IP. - Cache stable resources.
GET /v1/modelschanges rarely — cache it for minutes to hours instead of fetching it per request. - Watch
X-RateLimit-Remaining. If it trends toward zero, slow down before you hit the wall instead of after.