SDSignal Desk

Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads

Oct 5, 2026, 2:04 PM · MarkTechPost

Image: MarkTechPost

Reflection’s Beam doesn’t claim to be the strongest open model, it claims to be the most efficient one in its class, and the training details behind that pitch are the real story.

Why it matters

Beam is Reflection AI’s first open-weight model: a sparse mixture-of-experts with 501 billion total parameters and 23 billion active per token, aimed at coding, reasoning and agent work. Reflection says it competes with larger open models like GLM 5.2 while using three to four times less inference compute on reasoning benchmarks.

It isn’t self-hostable yet. Beam is in final red-teaming, early access runs through a waitlist, and Apache 2.0 weights are scheduled for later this month. Inference cost is what decides whether agents are affordable to run all day, so an efficiency-first frontier model is worth our attention.

From the desk

We appreciate the candor in the pitch. Reflection openly says Kimi K3 stays ahead on raw capability and is selling efficiency instead. That is a more useful claim for most buyers than another leaderboard crown, because agent workloads burn tokens constantly and a model that reaches a good answer with less compute changes the economics.

The reported numbers are strong but mixed. Reflection puts Beam at 80.9 on SWE-bench Verified, ahead of Nemotron 3 Ultra’s 70.7, and 80.1 on Terminal Bench v2.1, close to GLM 5.2 at 81.0 but behind DeepSeek V4.1 Flash and Kimi K3. These are Reflection’s own tables, with rival scores drawn from Artificial Analysis and DataCurve. Until outside testers run Beam on the released weights, we treat them as claims, not results.

The more interesting material is how it was built. Reinforcement learning is the central scaling axis here: about 10,500 GB300 GPUs for four weeks, more than 100 million rollouts, and roughly 1.3 billion sandboxes across nearly a million coding, agent and STEM environments. Reflection reports no plateau as RL compute increased, and says browsing skills improved even without browsing tasks in the mix. If that holds up, it’s evidence that heavy RL on agent tasks transfers across domains, which is good news for capability and a reason for safety teams to pay attention, since skills nobody trained for are also skills nobody specifically tested.

On that front, Reflection trained a separate safety and alignment teacher, merged it with the RL teacher through distillation, and used deliberative alignment. The safety evaluation results are promised in the technical report, not yet published. I’m watching for that before calling this a responsible release; for an open-weight model, there’s no recall once the weights are out.

There’s also a useful-AI angle in the controllable reasoning-effort setting and a length penalty that trained Beam to solve tasks with fewer tokens. Letting teams dial effort to task difficulty is exactly how agents get cheaper without getting dumber.

Context

Beam was pretrained on 23.8 trillion tokens in under four weeks on 6,144 GB300 GPUs, with midtraining extending effective context to 1 million tokens. It is text-only. Its load-balancing approach builds on auxiliary-loss-free balancing introduced with DeepSeek-V3.

Who feels it

Developers
A permissively licensed, efficiency-focused coding and agent model to evaluate once weights land, with a reasoning-effort dial for cost control.
Enterprises
Lower inference compute could make long-running agent workloads cheaper, pending independent benchmarks.
Safety researchers
Heavy RL with cross-domain skill transfer raises the importance of the promised safety evaluation data.

What to watch

  1. Publication of the Apache 2.0 weights and the full technical report, including safety evaluations
  2. Independent benchmark runs confirming the three-to-four-times compute efficiency claim
  3. Real-world agent performance at different reasoning-effort settings

Read the original

Continue at the source.

MarkTechPost

Also covering this