AERIOXFLUX
← AI Tools
AI Tools · open source

Reflection's Beam Is an American Open Model That Admits China Leads

A 501-billion-parameter open-weight model with 23 billion active, sold on inference efficiency rather than raw capability, with its own benchmark table showing Kimi K3 still ahead.

Flux Desk·2026-10-08·5 min read

Reflection AI has spent two years and roughly $4.7 billion of investor money, per PitchBook figures cited by TechCrunch, promising an American answer to China's open-weight models. On October 5 it showed the first piece of that answer. Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic work. The weights are not out yet. What is out is a long blog post and a benchmark table that, unusually for a launch, puts the competitors that beat it in plain view.

According to the company's blog, Beam is in final red-teaming and evaluation. Early access runs through a waitlist on Reflection's platform for a select group of users. The weights, a technical report, a model card and developer artifacts are due "later this month," with the weights under an Apache 2.0 license and tooling for running, evaluating and fine-tuning the model.

What Reflection built

The architecture section reads like a list of the stability tricks large open labs have converged on. Reflection describes 52 layers, interleaved local and global attention, fine-grained routed experts, and auxiliary-loss-free load balancing. The company says the final pretrained base has near-uniform expert use, with the busiest expert running at 1.04 times the average. Beam is text-only. Midtraining extends its effective context to 1 million tokens, a figure TechCrunch also reported, though reinforcement learning ran at a maximum of 256,000.

The training numbers are large and specific. Per the blog, Beam was pretrained on 23.8 trillion tokens of web, public and proprietary licensed data, in under four weeks on 6,144 Nvidia GB300 NVL72 GPUs, with goodput reaching 92.3% by the end of the run. Reflection reports nine semi-automatic rewinds, attributed to gradient spikes or suspected silent data corruption, and no large unrecoverable loss spikes.

Reinforcement learning was the bigger job. The blog says it ran on 10,500 GB300 GPUs over four weeks, produced more than 100 million rollouts, and used about 1.3 billion sandboxes across nearly a million environments, most of them synthetic. Reflection calls it "one of the largest scale RL runs conducted by any open lab to date" and says gains kept climbing with compute, with no plateau.

The table that shows who is ahead

The most telling part of the launch is the comparison table. Reflection benchmarks Beam against Moonshot's Kimi K3, Z.ai's GLM 5.3 and 5.2, Alibaba's Qwen 3.8 Max, DeepSeek V4.1 Flash, Nvidia's Nemotron 3 Ultra and Thinking Machines Lab's Inkling. On many rows, the Chinese models win.

On SWE-Bench Pro v2-Hard, Beam scores 77.2 against Kimi K3's 88.2 and GLM 5.3's 84.3. On Terminal-Bench v2.1, Beam's 80.1 trails GLM 5.3 at 88.2, Kimi K3 at 88.3 and DeepSeek V4.1 Flash at 90.6. On Humanity's Last Exam without tools, Beam posts 36.2 to Kimi K3's 46.9. On GPQA Diamond it scores 90.5 against Kimi K3's 93.5.

Where Beam wins, it mostly wins against Western peers. It scores 80.9 on SWE-Bench Verified to Inkling's 77.6 and Nemotron 3 Ultra's 70.7, and 78.0 on SWE-Bench Multilingual to Nemotron's 67.7. TechCrunch noted that on the four coding tests where both companies report results, Beam beats Inkling, which is multimodal where Beam is not. All of these are company-reported figures and none have been independently verified.

The pitch is efficiency

Reflection is not claiming the top of the open-weight leaderboard. Its claim is narrower. "Beam's advantage is efficiency at inference time," the blog says. The company says Beam is comparable to GLM-5.2 on advanced reasoning benchmarks while using roughly three to four times less inference compute, and that the gap widens against models above 2 trillion parameters such as Qwen 3.8-Max. TechCrunch put GLM-5.2 at about 744 billion total parameters with 40 billion active, which makes the comparison one of active-parameter count as much as training quality.

The methodology matters. Reflection estimates compute as twice the active parameters multiplied by mean generated tokens, excluding prefill and serving overhead. That is a reasonable first-order proxy, but it is the company's own yardstick, and buyers running long-context agent loops will care about the prefill it leaves out. A reasoning-effort setting lets users trade response length for accuracy.

For an enterprise, the trade is easy to state. A model that sits a few points behind Kimi K3 but costs a fraction as much to serve, with an Apache license and an American publisher, is a procurement argument. It is not a capability argument.

Who it is for

TechCrunch reported that Reflection calls Beam a "workhorse model" for enterprises, the public sector and developers. The company's commercial plan is what it calls AI factories, in which institutions train its models on their own data to build local systems. TechCrunch reported that Reflection is testing a sovereign AI factory partnership with South Korea's Shinsegae Group, and cited Axios reporting that hedge funds and trading firms are interested.

Distribution will run through hyperscalers and neoclouds, with integrations into open-source libraries at launch, per TechCrunch. The compute behind the next models is already contracted. TechCrunch reported that Reflection signed deals worth more than $7 billion combined with SpaceX and Nebius this summer, securing GB300 access through 2029. The blog says Beam is the first in a series and that its successors are already training.

What to watch

The weights. Until they ship, Beam is a blog post and a waitlist, and Reflection's numbers cannot be tested by anyone outside the company. The promised technical report and published safety evaluations will show how much of the efficiency claim survives independent measurement.

The second test is pricing once providers host it. If the three-to-four-times inference saving shows up in per-token rates, Beam has a real lane among enterprise and government buyers who want open weights without a Chinese publisher. If it does not, Reflection has shipped a credible model that, by its own table, still sits behind the open models it set out to answer.

#reflection-ai#beam#open-weights#mixture-of-experts#kimi-k3#glm

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.