Reflection has released Beam, a 501-billion-parameter open-weight Mixture-of-Experts model with 23 billion parameters active per token, built around coding, reasoning, and agentic workloads. Beam reaches competitive scores with leading open-weight coding and agentic models while using a fraction of the inference compute, making it a practical workhorse for enterprise deployments.
What is Beam and why does it matter?
Beam is a sparse Mixture-of-Experts model that activates only 23 billion of its 501 billion parameters for any given token. That ratio is the source of its efficiency story: it carries the capability of a frontier system without the inference cost of running one in full.
The model was pretrained on 23.8 trillion diverse, curated, and high-quality tokens drawn from the web and proprietary licensed datasets. According to the release, Beam matches or outperforms available similar-sized open base models at this stage of training.
How does Beam perform on coding, reasoning, and agentic benchmarks?
The release positions Beam as advancing the frontier for Western open-weight systems and as competitive with larger open models such as GLM 5.2, while approaching Qwen 3.8-Max on coding and agentic tasks. Frontier open models like Kimi K3 still lead on raw capability; Beam’s stated advantage is efficiency.
Selected results from the published benchmark table:
- DeepSWE v1.1: Beam 44.4, Qwen 3.8-Max 74.2, Kimi K3 68.0, GLM 5.2 44.0.
- SWE Bench Pro v2-Hard: Beam 77.2, GLM 5.3 84.3, Kimi K3 88.2.
- Terminal Bench v2.1: Beam 80.1, GLM 5.2 81.0, GLM 5.3 88.2, Kimi K3 88.3, Qwen 3.8-Max 88.3, DeepSeek V4.1 Flash 86.6, DeepSWE v1.1 90.6.
- SWE Bench Pro v1: Beam 65.5, Qwen 3.8-Max 67.7, GLM 5.2 62.1, Inkling 54.3.
- AIME 2026: Beam 97.8, GLM 5.2 99.2, Inkling 97.1.
- GPQA Diamond: Beam 90.5, GLM 5.2 91.2, GLM 5.3 91.7, Kimi K3 93.5, Qwen 3.8-Max 92.6, DeepSeek V4.1 Flash 90.9.
- HLE no tools: Beam 36.2, GLM 5.3 42.3, Kimi K3 46.9, Qwen 3.8-Max 43.6.
- AutomationBench public: Beam 37.0, GLM 5.3 48.2, Kimi K3 46.7, DeepSeek V4.1 Flash 54.8.
- MCP Atlas: Beam 78.7, GLM 5.3 84.2, Kimi K3 82.3, Qwen 3.8-Max 84.5.
- tau3 banking: Beam 38.0, Qwen 3.8-Max 55.2, GLM 5.2 37.1, Kimi K3 37.1.
- LongBench v2: Beam 65.5, GLM 5.2 64.0, Qwen 3.8-Max 66.3, Nemotron 3 Ultra 61.9.
- IFBench: Beam 79.7, Qwen 3.8-Max 82.8, Nemotron 3 Ultra 81.7, GLM 5.2 73.3.
The most important framing in the release is the compute efficiency angle: on advanced reasoning benchmarks, Beam achieves results comparable to GLM-5.2 while using roughly 3 to 4 times less inference compute. The gap is larger against models in the 2T+ parameter family like Qwen 3.8-Max, which require substantially more compute per generated token. Compute was estimated as FLOPs approximately equal to 2 times active parameter count times mean generated tokens per attempt, with multiply-add counted as two operations.
How was Beam trained at this scale?
High-compute reinforcement learning was treated as a central scaling axis for Beam. The team deployed 10.5 thousand NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes, and the team sourced about one million high-quality coding, agentic, and STEM environments to sustain it.
The release states that across the evaluation suite, capabilities continued to improve as RL compute increased, with no sign of a plateau. For context, Inkling was trained on 30 million rollouts and MiMo on 753 thousand.
Training used asynchronous policy gradients, a method that becomes harder to stabilize at scale as policy staleness grows and training-inference numerical mismatches compound. The team developed new algorithms to maintain stable learning under these conditions while systematically reducing training-inference mismatch. The result, per the release, is fully asynchronous RL that remains stable even when learning from interactions generated more than a day earlier: training was tested with one-day staleness at 107 weight versions behind the current policy and the numerics stayed stable.
Can Beam reason more efficiently on demand?
Beam was trained with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens. Early in RL, performance improved as completion lengths fell, meaning the model learned to solve tasks more effectively with less reasoning. Later, as agentic capabilities developed, completion lengths grew again, but those extra tokens supported further performance gains.
A reasoning effort parameter gives users a knob over how much thinking the model does. Lower settings favor shorter responses; higher settings allow longer reasoning for more demanding tasks. Across training, RL improved the trade-off between capability and token usage.
Does the training generalize beyond coding tasks?
During a training phase focused on reasoning, software engineering, and terminal tasks, Beam showed consistent gains in browsing despite the absence of browsing tasks from the RL mixture. The release frames this as evidence that Beam was learning broader agentic capabilities that generalize across domains. When given web access, the model organically learned to search for and query other large language models and to use OCR APIs to read documents.
Example tasks the release highlights:
- Building a live updating NYC subway map using publicly available NYC MTA data, including documentation lookups, authentication checks, geometry handling, and creative interpolation between stations.
- Creating a 3D astronaut free-fall game in p5.js where the astronaut dodges or destroys incoming asteroids, reasoned out in text even though Beam is text-only.
- Recreating a viral land versus water longitude-latitude grid puzzle with 16,200 points and getting 95.5% coverage right, sitting between Opus 5 at 92.5% and Fable 5 at 97.8%.
- Producing a fine-tuning notebook for the latest and smallest Gemma-4 model on a Text2SQL task, an out-of-distribution domain for Beam, raising Gemma’s accuracy on a held-out test set by 66.5%.
What infrastructure supported the RL run?
Beam’s training sustained an average of 110 thousand concurrent rollouts. Seven capabilities made that practical:
- Fully asynchronous execution: agents generate rollouts while the trainer learns and publishes new model versions, with version-aware staleness handling.
- Flexible compute allocation: inference-to-training GPU ratios shifted between 3.9:1 and 5.4:1, and the trainer was resized across five GPU mesh configurations without losing training state.
- Fast model updates: new weights reached the inference fleet in a median of about 12 seconds using hierarchical distribution via RoCE across racks and NVLink locally, cutting cross-rack traffic by 75% and making fleet-wide adoption 2.2 times faster.
- Resilience to inference failures: 71 inference incidents during the run were handled without terminating the training job, with a median recovery of eight minutes and lost capacity at 0.02% of elapsed serving GPU-minutes.
- Environments at scale: support for up to 170 thousand concurrent environments for the RL effort.
How does Beam relate to prior open-weight MoE work?
Beam is a sparse Mixture-of-Experts model, an architecture whose cost-effectiveness for very large models has been independently validated. DeepSeek-V3, a 671B total-parameter MoE with 37B activated per token, described in the DeepSeek-V3 technical report on arXiv, adopted Multi-head Latent Attention and DeepSeekMoE architectures, introduced an auxiliary-loss-free load balancing strategy, and trained on 14.8 trillion tokens. DeepSeek-V3 reported full training at 2.788M H800 GPU hours and made its checkpoints available. Beam operates in the same broad architectural family but with different parameter counts, a different active-per-token ratio, and a heavier reinforcement-learning phase, and the DeepSeek-V3 report is useful prior evidence that the MoE approach can deliver frontier-class efficiency at open-weight price tags.
How is Beam being released?
Beam is undergoing final red-teaming and evaluations. Weights, a technical report, a model card, and developer artifacts are scheduled for release later this month. Early access sign-up is available through the release.
FAQ
What is Reflection’s Beam model?
Beam is Reflection’s first open-weight model. It is a sparse Mixture-of-Experts model with 501 billion total parameters and 23 billion active parameters per token, designed for coding, reasoning, and agentic workloads.
How was Beam trained?
Beam was pretrained on 23.8 trillion diverse, curated, and high-quality tokens. Reinforcement learning used 10.5 thousand NVIDIA GB300 GPUs over four weeks, generating more than 100 million rollouts at up to 256K context length, with approximately 1.3 billion sandboxes and about one million training environments.
How efficient is Beam at inference?
According to the release, Beam achieves scores comparable to GLM-5.2 on advanced reasoning benchmarks while using 3 to 4 times less inference compute. The compute gap is larger against 2T+ parameter models like Qwen 3.8-Max, which require significantly more compute per generated token.
This article summarizes reporting from reflection.ai.

