On October 5, Reflection AI announced Beam, its first model with open weights. A sparse mixture-of-experts model, 501 billion parameters total, 23 billion active per token. Built for code, reasoning and agentic work.
Before this, there was nothing from Reflection to download. The company was founded in 2024 by former DeepMind researchers, I wrote about its $6.3 billion compute deal with SpaceX back in June, and in that same month it confirmed a funding round at a $25 billion pre-money valuation, with Nvidia, Sequoia and Citigroup on the cap table.
Pretraining, then a four-week RL run
Beam was pretrained on 23.8 trillion tokens from the web and from licensed datasets. Then came an unusually large RL run: over 100 million rollouts on 10,500 NVIDIA GB300s, for four weeks.
The benchmarks Reflection publishes itself: SWEBench Verified 80.9, SWE Bench Pro v1 65.5, Terminal Bench v2.1 80.1, MCP Atlas 78.7, GPQA Diamond 90.5, AIME 2026 97.8. On Humanity’s Last Exam without tools it scores 36.2 — Kimi K3 sits at 46.9 and Qwen 3.8-Max at 43.6 there, and Reflection says as much in its own post: on raw capability, the strongest open models stay ahead.
The pitch is elsewhere. On advanced reasoning, Reflection says Beam lands at GLM-5.2’s level while using three to four times less inference compute. Against models past two trillion parameters the gap widens further. More intelligence per token, cheaper to run — CEO Misha Laskin calls Beam a “workhorse” in his Semafor interview.
You can’t download anything yet. Beam is in final red-teaming and evaluation; weights, technical report, model card and developer artifacts are promised for later in October, with a waitlist in the meantime. The next model is already training, and Laskin says it will be considerably stronger.
The market Reflection is aiming at
Reflection isn’t selling Beam as a better model but as a Western one. Chinese labs have led open weights for years, and Western companies now reach for them because they’re cheaper on repetitive work. At the same time, US and UK evaluators found that recent Chinese models can help attackers exploit code vulnerabilities.
That’s exactly where Reflection’s offer sits: agencies, governments and enterprises building sovereign AI systems that don’t want to — or can’t — run Chinese models. “They don’t really have very good options today,” Laskin says. Thinking Machines Lab and Mistral play in the same space, and Axios reported on October 4 that several Western open models are due this month.
The number that matters isn’t 501
With 501 billion parameters in the headline, size is what you look at first. The three-to-four-times inference factor is the more interesting one. If you run an open model yourself, you don’t pay for parameters, you pay for tokens and GPU hours — and a model that reaches the same result on a quarter of the compute moves the figure that actually gets calculated.
None of it is checkable yet. The numbers come from Reflection, the comparisons from Artificial Analysis and DataCurve, and the weights aren’t anywhere. Until the download exists, Beam is a good story with a table next to it.