Rate limits
How fast each API may be called, how to read the waits and how to stay under.
RFS Router
Router limits are per credit account, not per key: more keys don't add room. A refused call is always free and never counts against the limits.
- Calls running at once
- 8 per account
- The next waits for one to finish
- Calls started
- 600 a minute per account
- Up to 60 at once after a quiet spell, then one every 100 ms
- Request body
- 4 MB
- Send images and files as URLs
- Streamed call
- 30 minutes
- Ends with a clean error event, charged for what was written
- Call that isn't streamed
- 13 minutes
- Stream anything long
- Smallest reserve
- $0.001
- Every call holds at least this while it runs
- Router keys
- 100 per account
- Revoke one to make room
- Key spending limit
- $0.01 to $1,000,000
- Per day, week or month, reset at 00:00 UTC
- Waits we name
- 1 to 60 s
- In Retry-After and Retry-After-Ms
8 calls running at once. The 9th is refused with 429 too_many_running_calls and a wait of 2 to 8 seconds, spread so a refused batch doesn't return all at once. A slot frees the moment a call ends.
600 calls started a minute. Starts are spaced on a schedule: one every 100 ms, with up to 60 at once after a quiet spell. Faster than that is refused with 429 rate_limited and the exact wait.
With calls that take a second or more, the 8-call limit is usually the one you meet first. The rate matters for many very short calls.
Reading the headers
Every 429 and 503 names its wait:
HTTP/1.1 429 Too Many Requests
retry-after: 2
retry-after-ms: 1180retry-after-msis exact;retry-afteris the same in whole seconds.- Waits run from 1 second to 60, with a little added at random so clients told the same wait don't all come back at once.
- The OpenAI and Anthropic SDKs read these headers and wait on their own.
Stay under
- Queue on your side. Run at most 8 calls per account at once. A simple semaphore does it.
- Ask for the output you need. A lower
max_tokensmeans a smaller reserve, so more calls fit in the balance at once. - Read what you ask for. A call whose client stops reading for 90 seconds is ended, and charged what the model wrote.
- Let the SDK retry. It already honors
Retry-After-Ms.
Management API
- Requests
- 300 a minute per secret key
- Past it: 429 rate_limited with Retry-After
- Launches and imports
- 10 a minute per project
- Across all of the project's keys
- Launches waiting to be sent
- 20 per project
- Prepared in the last two minutes and not sent
- Request body
- 3.5 MB
- A token image is at most 2 MB
- Idempotency-Key
- 24 hours
- A retry with the same key and body gets the first answer; a request that died mid-run frees its key after 6 minutes
- Page size
- 1 to 100
- 20 when not given
- Secret keys
- 10 per project
- Make them on the dashboard's Keys page
- Webhook endpoints
- 5 per project
- Each gets some or all events
- Fee strategies
- 20 in use per project
- Archive one to make room
Past a limit, a request is refused with 429 rate_limited and a Retry-After in seconds. A refused request takes nothing from any window, so retrying after the wait always fits.