Skip to main content
All /v1 endpoints share two independent per-minute rate limits: Both are fixed one-minute windows. Whichever limit you hit first applies. The defaults are well above normal create-and-poll traffic and may be tuned over time — read the response headers rather than hard-coding these numbers.
POST /v1/uploads has an additional, tighter cap of 60 uploads / minute per account — a rolling window — because it’s free and unmetered. It applies on top of the per-key and per-IP limits above and returns the same 429 + Retry-After. Regular generation and polling traffic isn’t affected.
Generation jobs also have a separate per-account concurrency cap — the number of jobs you can have in flight at once, shared with the web workshop. That cap is independent of the request limits on this page: polling status never counts against concurrency, and concurrency never shows up in the X-RateLimit-* headers.The cap is per account (not per key — all your keys and the web app share it), starts at your plan’s limit (up to 10 on Ultra), and rises automatically with sustained usage — heavy accounts are granted higher floors without asking, and we email you when it happens. Your current value is shown under Billing in the developer console, where you can also request an increase: tell us your use case and target concurrency, and the request is reviewed manually — you’ll get the decision by email, and an approved limit takes effect immediately.

Response headers

Every API response reports your current standing:

When you exceed a limit

Requests over the limit receive HTTP 429 with a Retry-After header (seconds) and this body:
Rejected requests still count against the window. A tight retry loop keeps the window full and you stay rate limited indefinitely. On a 429, stop and wait the full Retry-After before sending anything else.

Other responses that also return 429

Three further conditions answer with 429, and none of them is the per-minute limiter above. Read the body to tell them apart — only the rate limiter and the Grok cooldown send Retry-After.

Queue full

You have as many generations in flight as your plan’s concurrency cap allows. Waiting a fixed number of seconds is the wrong response: poll your running jobs and submit the next one as they finish. See the concurrency note above for the cap and how to raise it.

Monthly spend limit

Your account’s monthly API spend limit has been reached. It is a limit you set yourself, so retrying will not clear it — raise or remove it under Billing in the developer console.

Repeated Grok failures

Grok’s content filter rejects some prompts and images, and it is not always consistent about which. After three consecutive failed Grok generations within an hour, further Grok requests pause for ten minutes, growing by ten minutes with each additional failure; any success clears it. Other models are unaffected. Retrying the same prompt will not succeed — change the prompt or the input image, then retry after retry_after_seconds.

Staying under the limits

  • Poll gently. Poll job status every few seconds (images) to every 10–30 seconds (video) — not in a tight loop. See Async jobs.
  • Respect Retry-After. On 429, sleep for the advertised seconds, then resume. Add jitter if multiple workers share a key or IP.
  • Cache stable resources. GET /v1/models changes rarely — cache it for minutes to hours instead of fetching it per request.
  • Watch X-RateLimit-Remaining. If it trends toward zero, slow down before you hit the wall instead of after.