Developer tools

Build with Reflection Beam

Plan GPUs for self-hosting, estimate tokens and cost, generate working API code and set up your coding agent. Every number is tied to Reflection's own documentation.

Start with the code generator GPU memory planner

The essentials

Model ID
Beam-501B-A23B
Base URL
https://api.reflection.ai/openai/v1
Parameters
501B total, 23B active
Context window
262,144 tokens
Max output
131,072 tokens
Reasoning effort
low · medium · high · xhigh · max

Tools

Your first request in 30 seconds

The API follows OpenAI Chat Completions, so the official OpenAI SDKs work after you change the base URL and key.

bash
pip install openai
export REFLECTION_API_KEY="<your API key>"
python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.reflection.ai/openai/v1",
    api_key=os.environ["REFLECTION_API_KEY"],
)

messages = [
    {"role": "user", "content": "Explain the difference between a process and a thread in three sentences."},
]

completion = client.chat.completions.create(
    model="Beam-501B-A23B",
    messages=messages,
    reasoning_effort="medium",
)

message = completion.choices[0].message
print("Reasoning:", getattr(message, "reasoning_content", None))
print(message.content)

Six things that trip people up

Reasoning can't be turned off

Beam always reasons. reasoning_effort picks how much, from low to max, and medium is the default. Values like none or minimal return a 400 error.

Send reasoning back

In multi-turn chats and tool calls, return each earlier assistant message with its reasoning_content. Without it the API answers 400 missing_required_parameter.

Reasoning eats your token budget

Reasoning tokens count toward max_completion_tokens. Set it too low and you get finish_reason "length" with an empty answer.

Text only, Chat Completions only

No image, audio or file input, and no Responses, Embeddings or Batch APIs. Agents must use an OpenAI-compatible Chat Completions provider.

Some parameters do nothing

stop is accepted but ignored; n must be 1; logprobs and logit_bias are not supported; user and metadata return an error.

429 means wait, not retry now

Limits apply per organization, per minute and per UTC day. Honor Retry-After; x-should-retry: false means the daily budget is gone until 00:00 UTC.

More about Beam

FAQ

Can I download Beam's weights?

Not yet. Reflection has promised open weights under Apache 2.0 later in October 2026. Until then Beam is available only through the waitlisted beta API.

How much hardware does self-hosting need?

All 501 billion parameters must fit in memory: about 1 TB at BF16, about 500 GB at FP8 and about 250 GB at 4-bit, plus the KV cache. That is a multi-GPU server; use the memory planner for your setup.

Is the API compatible with the OpenAI SDK?

Yes for Chat Completions and Models. Point the SDK at https://api.reflection.ai/openai/v1 and use the model ID Beam-501B-A23B. Streaming, tool calling and JSON schema outputs work.

How much does the API cost?

Reflection has not published prices for the beta. The cost calculator lets you enter your own price and shows how reasoning tokens affect the total.