What is relayfor.si?

Rate limits

How fast each API may be called, how to read the waits and how to stay under.

RFS Router

Router limits are per credit account, not per key: more keys don't add room. A refused call is always free and never counts against the limits.

Calls running at once
8 per account
The next waits for one to finish
Calls started
600 a minute per account
Up to 60 at once after a quiet spell, then one every 100 ms
Request body
4 MB
Send images and files as URLs
Streamed call
30 minutes
Ends with a clean error event, charged for what was written
Call that isn't streamed
13 minutes
Stream anything long
Smallest reserve
$0.001
Every call holds at least this while it runs
Router keys
100 per account
Revoke one to make room
Key spending limit
$0.01 to $1,000,000
Per day, week or month, reset at 00:00 UTC
Waits we name
1 to 60 s
In Retry-After and Retry-After-Ms

8 calls running at once. The 9th is refused with 429 too_many_running_calls and a wait of 2 to 8 seconds, spread so a refused batch doesn't return all at once. A slot frees the moment a call ends.

600 calls started a minute. Starts are spaced on a schedule: one every 100 ms, with up to 60 at once after a quiet spell. Faster than that is refused with 429 rate_limited and the exact wait.

With calls that take a second or more, the 8-call limit is usually the one you meet first. The rate matters for many very short calls.

Reading the headers

Every 429 and 503 names its wait:

HTTP/1.1 429 Too Many Requests
retry-after: 2
retry-after-ms: 1180
  • retry-after-ms is exact; retry-after is the same in whole seconds.
  • Waits run from 1 second to 60, with a little added at random so clients told the same wait don't all come back at once.
  • The OpenAI and Anthropic SDKs read these headers and wait on their own.

Stay under

  • Queue on your side. Run at most 8 calls per account at once. A simple semaphore does it.
  • Ask for the output you need. A lower max_tokens means a smaller reserve, so more calls fit in the balance at once.
  • Read what you ask for. A call whose client stops reading for 90 seconds is ended, and charged what the model wrote.
  • Let the SDK retry. It already honors Retry-After-Ms.

Management API

Requests
300 a minute per secret key
Past it: 429 rate_limited with Retry-After
Launches and imports
10 a minute per project
Across all of the project's keys
Launches waiting to be sent
20 per project
Prepared in the last two minutes and not sent
Request body
3.5 MB
A token image is at most 2 MB
Idempotency-Key
24 hours
A retry with the same key and body gets the first answer; a request that died mid-run frees its key after 6 minutes
Page size
1 to 100
20 when not given
Secret keys
10 per project
Make them on the dashboard's Keys page
Webhook endpoints
5 per project
Each gets some or all events
Fee strategies
20 in use per project
Archive one to make room

Past a limit, a request is refused with 429 rate_limited and a Retry-After in seconds. A refused request takes nothing from any window, so retrying after the wait always fits.

On this page