RFS Router
Hundreds of models behind one address, in the OpenAI and Anthropic formats, paid by a token's credit.
The RFS Router is how your agents think. It serves hundreds of models, from Claude, GPT and Gemini to DeepSeek, Grok and Qwen, in the formats the official SDKs already speak. Each call is paid from the credit account of one token, through a router key.
One address
https://relayfor.si/api/v1The Anthropic SDKs take https://relayfor.si/api: they add /v1 themselves. Send a router key as Authorization: Bearer rf_ai_..., or as x-api-key.
| Format | Endpoint | Speaks it |
|---|---|---|
| OpenAI Chat Completions | POST /chat/completions | OpenAI SDKs, Vercel AI SDK, LangChain |
| OpenAI Responses | POST /responses | OpenAI SDKs |
| Anthropic Messages | POST /messages | Anthropic SDKs |
| Image generation | POST /images | Plain HTTP |
Every format works with every model: a Claude model on Chat Completions, a GPT model on Messages. GET /models lists them all with their prices, no key needed.
What a call costs
The model's own price plus 20%, charged to the key's account per call. Each answer says what it cost in usage.cost, in dollars, on every format but Messages. Nothing else is billed: no monthly fee, no seats, no minimum.
Before a call starts it reserves its worst case: the prompt plus the longest answer max_tokens allows. You pay only what the model wrote, and the rest goes back the moment it ends. How a call is paid
Always set max_tokens
Without a limit, a call reserves the model's whole output, which can be dollars, and a
small balance refuses it with 402 insufficient_balance. A limit costs nothing: you
pay only what the model writes.
What it keeps
Nothing you send. Prompts and answers are never stored, and nothing carries over between calls: send the conversation every time. For each call relayfor.si keeps only its model, its tokens, its cost, its key and its time, for your usage and ledger.
Limits
An account runs at most 8 calls at once and starts at most 600 a minute. A refused call is free, and its answer names the wait. See Rate limits.