# Reflection Beam: 501B Open-Weight Model, Benchmarks

> Reflection AI's Beam is a 501B open-weight MoE with 23B active, built for coding and agents. Specs, the benchmark table, and when the weights land.

- Author: Nitish Garg, Founder & CEO, CellCog
- Published: 2026-10-06
- Canonical (HTML): https://cellcog.ai/blog/reflection-beam-open-weight-model/
- Section: Guides / Choosing a platform
- Publisher: CellCog (https://cellcog.ai), the AI employee platform. Blog index for agents: https://cellcog.ai/blog/llms.txt

## Key points

- Reflection AI announced Beam on October 5, 2026, its first open-weight model: a sparse mixture-of-experts with 501 billion total parameters and 23 billion active.
- The weights are not out yet. Reflection says the weights, a technical report and a model card come later in October under an Apache 2.0 license; early access is by sign-up now.
- Beam is text-only and aimed at coding, reasoning and agent work; TechCrunch reports a 1 million token context window.
- Reflection says Beam scores like GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute, by its own estimate.
- Its reinforcement learning run produced more than 100 million rollouts on 10,500 Nvidia GB300 GPUs over four weeks.
- On Reflection's own table, Beam trails GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most reported rows, and no outside lab has verified the scores yet.

## At a glance

- **What is Reflection Beam?** Reflection AI's first open-weight model: a text-only mixture-of-experts with 501B total and 23B active parameters, built for coding and agents.
- **Can I download it?** Not yet. Reflection says the weights, technical report and model card come later in October 2026, under Apache 2.0; early access is by sign-up.
- **How good is it?** On Reflection's own table it is close to GLM 5.2 and behind GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most rows; its pitch is lower inference compute.

**Reflection AI, the Brooklyn lab founded in 2024 by two former Google DeepMind researchers, announced Beam on October 5, 2026: its first open-weight model, a sparse mixture-of-experts with 501 billion total parameters and 23 billion active, built for coding, reasoning and agent work.** The weights are not out yet. Reflection says Beam is in final red-teaming and evaluations, with early access by sign-up, and "We will release the weights, technical report, model card, and developer artifacts later this month." This page reads [Reflection's announcement](https://reflection.ai/blog/introducing-beam), its [post on X](https://x.com/reflection_ai/status/2107186849370247235) (19:11 UTC on October 5) and [TechCrunch's report](https://techcrunch.com/2026/10/05/reflection-debuts-beam-a-open-weight-ai-model-to-rival-chinese-models-at-lower-compute-cost/), as of October 6, 2026.

## What Reflection announced

*Table: Reflection Beam at announcement (Reflection blog and TechCrunch, read October 6, 2026)*

| Item | Beam |
|---|---|
| Architecture | Sparse mixture-of-experts, text only |
| Total parameters | 501 billion |
| Active parameters | 23 billion |
| Pretraining data | 23.8 trillion tokens |
| Context window | 1 million tokens (per TechCrunch) |
| Reinforcement learning | 100M+ rollouts on 10,500 Nvidia GB300 GPUs over 4 weeks |
| License | Apache 2.0, at weights release (per Reflection) |
| Weights, tech report, model card | Later in October 2026 |
| Access on October 5 | Early access sign-up |

Reflection positions Beam as a Western answer to the Chinese open-weight models. In its own words, Beam "is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks," while it concedes the top of the table: "Where frontier open models like Kimi K3 remain ahead on raw capability," its case is efficiency at inference time.

## The benchmarks, from Reflection's own table

Every number below comes from Reflection's announcement. Reflection says it took the other models' results from Artificial Analysis and DataCurve; NR means no score was reported. TechCrunch notes the claims have not been independently verified.

*Table: Selected rows from Reflection's benchmark table (Reflection, October 5, 2026; NR = not reported)*

| Benchmark | Beam | Inkling | GLM 5.2 | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|---|
| Terminal Bench v2.1 | 80.1 | 63.8 | 81.0 | 88.2 | 88.3 | 86.6 | 90.6 |
| DeepSWE v1.1 | 44.4 | NR | 44.0 | 61.0 | 68.0 | 51.0 | 74.2 |
| SWE Bench Pro v1 | 65.5 | 54.3 | 62.1 | NR | NR | 67.7 | NR |
| HLE, no tools | 36.2 | 29.7 | 40.5 | 42.3 | 46.9 | 43.6 | 39.1 |
| GPQA Diamond | 90.5 | 87.2 | 91.2 | 91.7 | 93.5 | 92.6 | 90.9 |
| MCP Atlas | 78.7 | 76.0 | 77.8 | 84.2 | 82.3 | 84.5 | NR |
| AutomationBench (public) | 37.0 | NR | 26.2 | 48.2 | 46.7 | 39.8 | 54.8 |

Read plainly, the table puts Beam level with GLM 5.2 and ahead of Inkling, Thinking Machines Lab's open model, on the rows where both report. It sits behind GLM 5.3, Kimi K3 and DeepSeek V4.1 Flash on most rows. Inkling is multimodal and Beam is text-only, as TechCrunch points out.

Reflection's real claim is about cost. It says Beam reaches scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute. That figure is an estimate: Reflection computes compute as roughly two times active parameters times generated tokens, and says it leaves out prompt prefill, attention costs and serving overhead, so it is "an approximate compute comparison rather than measured inference cost."

## How it was trained

Reflection made reinforcement learning the main scaling lever. The RL run used 10,500 Nvidia GB300 GPUs for four weeks, generated more than 100 million rollouts with up to 256,000 tokens of context, used about 1.3 billion sandboxes for training and grading, and drew on one million coding, agent and STEM environments. "We believe this is one of the largest scale RL runs conducted by any open lab to date."

Two details stand out for anyone who runs agents. Reflection says training stayed stable even when learning from rollouts generated more than a day earlier, up to 107 weight versions behind the current policy. And Beam picked up browsing skills it was never trained on: during a phase of reasoning, software and terminal tasks, its browsing scores rose anyway, and when given web access it learned on its own to search for and query other language models and to call OCR services to read documents. Users trade speed for depth with a reasoning effort setting.

## Who it is for

Reflection pitches Beam to enterprises, the public sector and developers as a workhorse model. TechCrunch adds the business context: Reflection has raised roughly $4.7 billion, per PitchBook, at a $25 billion pre-money valuation in its last round; it signed compute deals worth more than $7 billion with SpaceX and Nebius for Nvidia GB300 chips through 2029; and it is testing a sovereign AI factory partnership with Shinsegae Group in South Korea. At launch, Reflection says Beam will be distributed through hyperscalers and neoclouds.

## What this means if you run agents

A model like Beam matters in two ways. For teams that must keep data and models on their own hardware, a 501B open-weight model that is strong at coding and tool use widens the choice beyond Chinese labs. For everyone else, open models set the floor on price that closed APIs are measured against. One practical note: a mixture-of-experts model still has to hold all of its weights in memory to serve, so 23 billion active parameters cut compute per token, not the hardware needed to load it.

CellCog's AI employees run on Claude Opus 5.5 at every tier, not on Beam. We will read the technical report and outside evaluations when they arrive.

## What we are watching

- **The weights, technical report and model card**, which Reflection says arrive later in October.
- **Independent evaluations** of the benchmark and compute claims.
- **The Apache 2.0 release**, and whether the weights arrive on that license as promised.

## Sources

- Reflection AI, ["Introducing Beam: Reflection's 501B open-weight model"](https://reflection.ai/blog/introducing-beam), October 5, 2026
- Reflection AI on X, [announcement post](https://x.com/reflection_ai/status/2107186849370247235), October 5, 2026
- TechCrunch, ["Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost"](https://techcrunch.com/2026/10/05/reflection-debuts-beam-a-open-weight-ai-model-to-rival-chinese-models-at-lower-compute-cost/), October 5, 2026

## FAQ

**Is Reflection Beam open source?**

It is announced as an open-weight model, but on October 5, 2026 the weights were not downloadable. Reflection says it will release the weights, a technical report, a model card and developer artifacts later in October, after final red-teaming and evaluations, and says the weights will be under an Apache 2.0 license.

**How big is Beam?**

501 billion total parameters with 23 billion active per token, pretrained on 23.8 trillion tokens, per Reflection. TechCrunch reports a 1 million token context window. For scale, TechCrunch puts GLM-5.2 at roughly 744 billion total and 40 billion active.

**How does Beam compare with GLM 5.3 and Kimi K3?**

On Reflection's own benchmark table, Beam is behind both on most rows where they report scores, for example 80.1 on Terminal Bench v2.1 against 88.2 for GLM 5.3 and 88.3 for Kimi K3. Reflection's case is efficiency: similar reasoning scores to GLM-5.2 for 3 to 4 times less inference compute, by its own estimate.

**Can I try Beam now?**

Only through Reflection's early access sign-up. TechCrunch reports that at launch Beam will be distributed through hyperscalers and neoclouds, with integrations across open source libraries.

**Does CellCog use Beam?**

No. CellCog's AI employees run on Claude Opus 5.5 at every tier. We track open-weight models like Beam because they set the floor on price and are the option for teams that must run models on their own hardware.

## Related

- [GLM 5.3 for AI Agents: Release Date, Weights, What It Means](https://cellcog.ai/blog/glm-5-3-for-ai-agents/index.md)
- [Kolibri: Aleph Alpha's 78B Open Model, Benchmarked](https://cellcog.ai/blog/aleph-alpha-kolibri/index.md)
- [GLM 5.3 vs Qwen3.8-Max: The Open-Weight Frontier, Compared (August 2026)](https://cellcog.ai/blog/glm-5-3-vs-qwen3-8-max/index.md)
- [Best AI Agent Harnesses: October 2026 Rankings](https://cellcog.ai/blog/best-ai-agent-harnesses/index.md)

## The AI employee for this read

[AI Software Engineer](https://cellcog.ai/ai-employees/ai-software-engineer): I built this page. For what it covers, hire an engineer: it works in your repo behind an approval gate, so nothing reaches your world unclassified.

---

Markdown alternate of https://cellcog.ai/blog/reflection-beam-open-weight-model/. Try CellCog free, no credit card needed: https://cellcog.ai/signup
