Reasoning can't be turned off
Beam always reasons. reasoning_effort picks how much, from low to max, and medium is the default. Values like none or minimal return a 400 error.
Developer tools
Plan GPUs for self-hosting, estimate tokens and cost, generate working API code and set up your coding agent. Every number is tied to Reflection's own documentation.
How much VRAM Beam needs at each precision and context length, and how many GPUs that is.
02 / costEstimate tokens per request with reasoning, daily volume against rate limits, and self-hosting cost.
03 / codeWorking curl, Python and TypeScript for streaming, tool calls, JSON output and multi-turn chats.
04 / agentsCopy-ready setup for Mirror CLI, Pi, OpenCode and Hermes Agent with Beam.
The API follows OpenAI Chat Completions, so the official OpenAI SDKs work after you change the base URL and key.
pip install openai
export REFLECTION_API_KEY="<your API key>" import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.reflection.ai/openai/v1",
api_key=os.environ["REFLECTION_API_KEY"],
)
messages = [
{"role": "user", "content": "Explain the difference between a process and a thread in three sentences."},
]
completion = client.chat.completions.create(
model="Beam-501B-A23B",
messages=messages,
reasoning_effort="medium",
)
message = completion.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print(message.content) Beam always reasons. reasoning_effort picks how much, from low to max, and medium is the default. Values like none or minimal return a 400 error.
In multi-turn chats and tool calls, return each earlier assistant message with its reasoning_content. Without it the API answers 400 missing_required_parameter.
Reasoning tokens count toward max_completion_tokens. Set it too low and you get finish_reason "length" with an empty answer.
No image, audio or file input, and no Responses, Embeddings or Batch APIs. Agents must use an OpenAI-compatible Chat Completions provider.
stop is accepted but ignored; n must be 1; logprobs and logit_bias are not supported; user and metadata return an error.
Limits apply per organization, per minute and per UTC day. Honor Retry-After; x-should-retry: false means the daily budget is gone until 00:00 UTC.
Not yet. Reflection has promised open weights under Apache 2.0 later in October 2026. Until then Beam is available only through the waitlisted beta API.
All 501 billion parameters must fit in memory: about 1 TB at BF16, about 500 GB at FP8 and about 250 GB at 4-bit, plus the KV cache. That is a multi-GPU server; use the memory planner for your setup.
Yes for Chat Completions and Models. Point the SDK at https://api.reflection.ai/openai/v1 and use the model ID Beam-501B-A23B. Streaming, tool calling and JSON schema outputs work.
Reflection has not published prices for the beta. The cost calculator lets you enter your own price and shows how reasoning tokens affect the total.