Skip to content
AI EmployeeSuper-AgentsAgent-to-AgentTutorialsPricingBlogContact

Reflection Beam: 501B Open-Weight Model, Benchmarks

At a glanceQuick answers
What is Reflection Beam?
Reflection AI’s first open-weight model: a text-only mixture-of-experts with 501B total and 23B active parameters, built for coding and agents.
Can I download it?
Not yet. Reflection says the weights, technical report and model card come later in October 2026, under Apache 2.0; early access is by sign-up.
How good is it?
On Reflection’s own table it is close to GLM 5.2 and behind GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most rows; its pitch is lower inference compute.
Data illustration on off-white paper: a wide teal beam enters a glass prism and a thin amber beam leaves it, beside the large figure 501B total parameters and 23B active in amber
Fig 0A huge model with a narrow working beam: 501 billion parameters, 23 billion active. Made by CellCog's image agent, running GPT Image 2.5.

Reflection AI, the Brooklyn lab founded in 2024 by two former Google DeepMind researchers, announced Beam on October 5, 2026: its first open-weight model, a sparse mixture-of-experts with 501 billion total parameters and 23 billion active, built for coding, reasoning and agent work. The weights are not out yet. Reflection says Beam is in final red-teaming and evaluations, with early access by sign-up, and “We will release the weights, technical report, model card, and developer artifacts later this month.” This page reads Reflection’s announcement, its post on X (19:11 UTC on October 5) and TechCrunch’s report, as of October 6, 2026.

On this page · 7 sectionsOpen
  1. What Reflection announced
  2. The benchmarks, from Reflection’s own table
  3. How it was trained
  4. Who it is for
  5. What this means if you run agents
  6. What we are watching
  7. Sources
Key points6 · 6 min full read
  1. A grid of teal dots with only three lit amber: a large model with few parts active at once.
    Reflection AI announced Beam on October 5, 2026, its first open-weight model: a sparse mixture-of-experts with 501 billion total parameters and 23 billion active.
  2. An open padlock, a download arrow and a calendar: weights promised for later in the month.
    The weights are not out yet. Reflection says the weights, a technical report and a model card come later in October under an Apache 2.0 license; early access is by sign-up now.
  3. A code window beside a terminal prompt: coding and agent tasks.
    Beam is text-only and aimed at coding, reasoning and agent work; TechCrunch reports a 1 million token context window.
  4. A speed gauge with the needle low beside a coin: less compute per answer.
    Reflection says Beam scores like GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute, by its own estimate.
  5. Three server racks inside a circular arrow: a long training run on GPUs.
    Its reinforcement learning run produced more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks.
  6. A short bar with an amber cap beside two taller bars: still behind the leading open models.
    On Reflection’s own table, Beam trails GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most reported rows, and no outside lab has verified the scores yet.

§ 01What Reflection announced

Item Beam
Architecture Sparse mixture-of-experts, text only
Total parameters 501 billion
Active parameters 23 billion
Pretraining data 23.8 trillion tokens
Context window 1 million tokens (per TechCrunch)
Reinforcement learning 100M+ rollouts on 10,500 Nvidia GB300 GPUs over 4 weeks
License Apache 2.0, at weights release (per Reflection)
Weights, tech report, model card Later in October 2026
Access on October 5 Early access sign-up
Table 1Reflection Beam at announcement (Reflection blog and TechCrunch, read October 6, 2026)

Reflection positions Beam as a Western answer to the Chinese open-weight models. In its own words, Beam “is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks,” while it concedes the top of the table: “Where frontier open models like Kimi K3 remain ahead on raw capability,” its case is efficiency at inference time.

§ 02The benchmarks, from Reflection’s own table

Every number below comes from Reflection’s announcement. Reflection says it took the other models’ results from Artificial Analysis and DataCurve; NR means no score was reported. TechCrunch notes the claims have not been independently verified.

Benchmark Beam Inkling GLM 5.2 GLM 5.3 Kimi K3 Qwen 3.8 Max DeepSeek V4.1 Flash
Terminal Bench v2.1 80.1 63.8 81.0 88.2 88.3 86.6 90.6
DeepSWE v1.1 44.4 NR 44.0 61.0 68.0 51.0 74.2
SWE Bench Pro v1 65.5 54.3 62.1 NR NR 67.7 NR
HLE, no tools 36.2 29.7 40.5 42.3 46.9 43.6 39.1
GPQA Diamond 90.5 87.2 91.2 91.7 93.5 92.6 90.9
MCP Atlas 78.7 76.0 77.8 84.2 82.3 84.5 NR
AutomationBench (public) 37.0 NR 26.2 48.2 46.7 39.8 54.8
Scroll to compare all columns
Table 2Selected rows from Reflection’s benchmark table (Reflection, October 5, 2026; NR = not reported)
Terminal Bench v2.1 scores, from Reflection's own tableBar chart of Terminal Bench v2.1 scores from Reflection's table: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM 5.3 88.2, Qwen 3.8 Max 86.6, GLM 5.2 81.0, Beam 80.1 highlighted, Inkling 63.8DeepSeek V4.1 Flash90.6Kimi K388.3GLM 5.388.2Qwen 3.8 Max86.6GLM 5.281.0Beam80.1Inkling63.8Terminal Bench v2.1 scores, from Reflection's own tableBar chart of Terminal Bench v2.1 scores from Reflection's table: DeepSeek V4.1 Flash 90.6, Kimi K3 88.3, GLM 5.3 88.2, Qwen 3.8 Max 86.6, GLM 5.2 81.0, Beam 80.1 highlighted, Inkling 63.8DeepSeek V4.1 Flash90.6Kimi K388.3GLM 5.388.2Qwen 3.8 Max86.6GLM 5.281.0Beam80.1Inkling63.8
Fig 1Terminal Bench v2.1 scores, from Reflection's own table

Read plainly, the table puts Beam level with GLM 5.2 and ahead of Inkling, Thinking Machines Lab’s open model, on the rows where both report. It sits behind GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most rows. Inkling is multimodal and Beam is text-only, as TechCrunch points out.

Reflection’s real claim is about cost. It says Beam reaches scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute. That figure is an estimate: Reflection computes compute as roughly two times active parameters times generated tokens, and says it leaves out prompt prefill, attention costs and serving overhead, so it is “an approximate compute comparison rather than measured inference cost.”

§ 03How it was trained

Reflection made reinforcement learning the main scaling lever. The RL run used 10,500 Nvidia GB300 GPUs for four weeks, generated more than 100 million rollouts with up to 256,000 tokens of context, used about 1.3 billion sandboxes for training and grading, and drew on one million coding, agent and STEM environments. “We believe this is one of the largest scale RL runs conducted by any open lab to date.”

Two details stand out for anyone who runs agents. Reflection says training stayed stable even when learning from rollouts generated more than a day earlier, up to 107 weight versions behind the current policy. And Beam picked up browsing skills it was never trained on: during a phase of reasoning, software and terminal tasks, its browsing scores rose anyway, and when given web access it learned on its own to search for and query other language models and to call OCR services to read documents. Users trade speed for depth with a reasoning effort setting.

§ 04Who it is for

Reflection pitches Beam to enterprises, the public sector and developers as a workhorse model. TechCrunch adds the business context: Reflection has raised roughly $4.7 billion, per PitchBook, at a $25 billion pre-money valuation in its last round; it signed compute deals worth more than $7 billion with SpaceX and Nebius for Nvidia GB300 chips through 2029; and it is testing a sovereign AI factory partnership with Shinsegae Group in South Korea. At launch, Reflection says Beam will be distributed through hyperscalers and neoclouds.

§ 05What this means if you run agents

A model like Beam matters in two ways. For teams that must keep data and models on their own hardware, a 501B open-weight model that is strong at coding and tool use widens the choice beyond Chinese labs. For everyone else, open models set the floor on price that closed APIs are measured against. One practical note: a mixture-of-experts model still has to hold all of its weights in memory to serve, so 23 billion active parameters cut compute per token, not the hardware needed to load it.

CellCog’s AI employees run on Claude Opus 5.5 at every tier, not on Beam. We will read the technical report and outside evaluations when they arrive.

§ 06What we are watching

  • The weights, technical report and model card, which Reflection says arrive later in October.
  • Independent evaluations of the benchmark and compute claims.
  • The Apache 2.0 release, and whether the weights arrive on that license as promised.

§ 07Sources

Frequently asked5 questions

Q1Is Reflection Beam open source?

It is announced as an open-weight model, but on October 5, 2026 the weights were not downloadable. Reflection says it will release the weights, a technical report, a model card and developer artifacts later in October, after final red-teaming and evaluations, and says the weights will be under an Apache 2.0 license.

Q2How big is Beam?

501 billion total parameters with 23 billion active per token, pretrained on 23.8 trillion tokens, per Reflection. TechCrunch reports a 1 million token context window. For scale, TechCrunch puts GLM-5.2 at roughly 744 billion total and 40 billion active.

Q3How does Beam compare with GLM 5.3 and Kimi K3?

On Reflection’s own benchmark table, Beam is behind both on most rows where they report scores, for example 80.1 on Terminal Bench v2.1 against 88.2 for GLM 5.3 and 88.3 for Kimi K3. Reflection’s case is efficiency: similar reasoning scores to GLM-5.2 for 3 to 4 times less inference compute, by its own estimate.

Q4Can I try Beam now?

Only through Reflection’s early access sign-up. TechCrunch reports that at launch Beam will be distributed through hyperscalers and neoclouds, with integrations across open source libraries.

Q5Does CellCog use Beam?

No. CellCog’s AI employees run on Claude Opus 5.5 at every tier. We track open-weight models like Beam because they set the floor on price and are the option for teams that must run models on their own hardware.

Published 06 October 2026 All Choosing a platform →