What is relayfor.si?

Reasoning

Let a model think before it answers, and what thinking costs.

Reasoning models think before they answer. You control how much.

{
  "model": "openai/gpt-6.1-sol",
  "reasoning": { "effort": "medium" },
  "max_tokens": 4096
}

Which levels a model takes

Each model takes its own few levels. GET /models lists them under reasoning:

{
  "id": "z-ai/glm-5.3",
  "reasoning": {
    "supported_efforts": ["max", "high", "low"],
    "default_effort": "max",
    "mandatory": true,
    "supports_max_tokens": false
  }
}
  • supported_efforts: the effort values it takes, deepest first. Empty when it thinks the same way every time.
  • default_effort: how it thinks when the request names no effort.
  • mandatory: it always thinks, so "none" is never listed for it.
  • supports_max_tokens: it thinks within a share of max_tokens, so give it room to answer.

reasoning is null for a model that does not reason.

What it costs

Thinking is billed as output, at the model's output price. It counts toward max_tokens, so leave room for both the thinking and the answer.

  • effort: "low" answers faster and costs less. Good for simple agent steps.
  • effort: "high" thinks longest. Save it for hard problems.

Streams send the thinking as it happens, before the answer. In the playground it shows in a Thinking fold above the answer.

On this page