Beam Does Not Take Down DeepSeek, Kimi, or GLM: Reflection's Own Table Puts It Behind All Three
Watch on TikTok
A 50.478-second clip (0:50) at 720x1280, H.264 High profile at 30 fps, AAC HE-AACv2 stereo at 44.1 kHz, 5,558,507 bytes on disk (5.3 MB, 881 kbps overall, 810 kbps video, 64 kbps audio), posted 2026-10-05 at 22:46:57 UTC by @thesanj23, channel nickname "Sanjay Ravi," carrying 328 views, 19 likes, 4 comments, and 1 repost when captured on 2026-10-06. The audio track is listed as "original sound - Sanjay Ravi" by Sanjay Ravi, which is the speaker's own voice with no music bed. I read 18 of the 25 sampled frames plus three upscaled stills pulled straight from the mp4 at t=3.0s, t=40.0s, and t=41.0s, against a 194-word transcript. The footage is a single static selfie shot: a man in a navy crewneck sweatshirt in a mesh-backed office chair, cream wall behind him, open white-framed doorway over his right shoulder, a smoke detector on the ceiling, and dried flowers at the left edge. Burned-in karaoke captions run across his chest in white with the stressed word in yellow, stepping through "AI COMPANY," "REFLECTION FOR," "OF 2025," "WHAT'S EXCITING," "CHINA HAS," "A LOT OF MY," "AND SERVE," "MODELS COMING," "AI APPLICATIONS," "A DUMB MODEL," "IT'S UP THERE," "AS FAR," "NEMOTRON," "YOU GUYS FOLLOW," and "COOL." Two screen recordings cut in over the face. The first, roughly 0:02 to 0:06, is a green panel from Reflection's own site with a pill-shaped nav bar reading "Reflection" next to a two-circle logomark, five decorative circular images (a rocket launch, an airliner, a green cell-like texture, a pale blue swirl, a striped disc), the headline "Introducing Beam: Reflection's 501B open-weight model," and the subhead "We build open models for everyone to access, use, and build on." The second, around 0:39 to 0:41, is Reflection's interactive benchmark table. A green pill tab reading "AGENTIC CODING/TERMINAL" is selected; greyed tabs for "REASONING" and a truncated "TC" sit to its right with a chevron indicating more columns off-screen. Only two of the eight result columns fit the vertical crop: "Beam" and "Inkling." The visible rows are DeepSWE v1.1 at 44.4 and NR, SWE Bench Pro v2-Hard at 77.2 and 56.9, SWE Bench Pro v1 at 65.5 and 54.3, Terminal Bench v2.1 at 80.1 and 63.8, SWE Atlas Codebase QnA at 34.6 and NR, SWEBench Multilingual at 78.0 and NR, and SWEBench Verified at 80.9 and 77.6. The six comparison columns the crop hides are exactly the ones that decide whether the video's headline holds.
The headline claim is contradicted by the table on screen
The video's own title asks whether Beam will "take down DeepSeek, Kimi, and GLM," and the transcript softens it to "based on the benchmarks, it's up there with GLM and Kimi and DeepSeq as far as intelligence and agentic workflows." Reflection's Beam announcement publishes the full table that the clip crops. On the same agentic coding and terminal tab the video shows, Beam scores 44.4 on DeepSWE v1.1 against DeepSeek V4.1 at 74.2, Kimi K3 at 68.0, and GLM 5.3 at 61.0. On Terminal Bench v2.1, Beam's 80.1 sits below GLM 5.2 at 81.0, GLM 5.3 at 88.2, Kimi K3 at 88.3, and DeepSeek V4.1 at 90.6. On SWE Bench Pro v2-Hard, Beam's 77.2 trails GLM 5.3 at 84.3 and Kimi K3 at 88.2. The one coding row where Beam leads a Chinese model outright is SWE Bench Pro v1, where its 65.5 beats GLM 5.2's 62.1 while still losing to Qwen 3.8 Max at 67.7.
The reasoning tab tells the same story. Beam's 36.2 on HLE with no tools sits below GLM 5.2 at 40.5, GLM 5.3 at 42.3, Kimi K3 at 46.9, Qwen 3.8 Max at 43.6, and DeepSeek V4.1 at 39.1. On GPQA Diamond, Beam's 90.5 is the lowest of the six models that reported a score except Nemotron 3 Ultra. Reflection does not claim otherwise. Its own framing is that Beam "achieves scores comparable to GLM-5.2 while using 3-4x less inference compute," which is a claim about the previous GLM generation and about efficiency, not a claim of leadership.
"Released" is premature: the weights are not out yet
The transcript says "today Reflection released Beam, which is their first open weight LLM." The first half of that is wrong on 2026-10-05 and still wrong on 2026-10-06. Reflection's post states that it "will release the weights, technical report, model card, and developer artifacts later this month," with the model "undergoing final red-teaming and evaluations" and an early-access waitlist in place. The announced license is Apache 2.0. What shipped on announcement day was a blog post, a benchmark table, and a signup form.
The second half of the sentence is correct. Beam is Reflection's first open-weight model, a text-only sparse mixture-of-experts system with 501 billion total parameters and 23 billion active per token, a one-million-token context window, pretrained on 23.8 trillion tokens. Reflection reports pretraining on 6,144 NVIDIA GB300 NVL72 GPUs in under four weeks and a reinforcement-learning run on 10,500 GB300s for four weeks producing more than 100 million rollouts at up to 256K context.
Every number in the video is self-reported and unreplicated
Nothing in the on-screen table has been reproduced by a third party. Reflection ran the evaluations, chose the comparison set, chose which cells to fill and which to mark NR, and published the result alongside a product announcement. That is standard practice for a model launch and it is also the reason the figures should carry a label.
The efficiency claim needs a sharper caveat than the benchmark scores do. Reflection's own methodology note reads: "Generated tokens include both reasoning and the final answer. For mixture-of-experts models, we used the parameters activated per token rather than the total model size. These estimates exclude prompt prefill, context-dependent attention operations, and serving overhead, so they represent an approximate compute comparison rather than measured inference cost." The 3x to 4x figure is arithmetic on active-parameter counts and generated-token counts, not a measurement of what serving Beam costs. Prefill and attention are exactly where long-context agentic workloads spend their money, and both are excluded.
The comparison is also not like-for-like across rows. Several cells are marked NR because the competing lab never published that benchmark, so the table's coverage varies by column rather than reflecting a uniform sweep. Scores produced by different labs on agentic benchmarks depend heavily on the scaffold, the tool harness, the retry budget, and the subset of tasks run. Terminal Bench and SWE-bench results in particular move several points on harness choice alone.
The Nemotron comparison is right on the headline rows and wrong on others
The transcript's last benchmark claim is "it's actually more intelligent than Nvidia's NemoTron Ultra model," with the caption reading "NEMOTRON." The model is NVIDIA's Nemotron 3 Ultra, a 550-billion-parameter open-weight model with roughly 55 billion active per token on a hybrid Mamba-attention MoE architecture, announced at Computex on June 4, 2026 under a permissive Linux Foundation license with weights, training data, and recipes published.
On Reflection's table, Beam leads Nemotron 3 Ultra on most rows: SWE Bench Pro v1 at 65.5 versus 46.4, Terminal Bench v2.1 at 80.1 versus 56.4, SWEBench Verified at 80.9 versus 70.7, SWEBench Multilingual at 78.0 versus 67.7, HLE at 36.2 versus 26.7, GPQA Diamond at 90.5 versus 87.0, MCP Atlas at 78.7 versus 63.1, and BrowseComp at 77.4 versus 44.4. It does not sweep. Nemotron 3 Ultra leads on IFBench at 81.7 versus 79.7, ties on AA-LCR at 79.3, and both trail on AA Omniscience where Beam's 13.0 beats Nemotron's 8.6. Calling Beam "more intelligent" is defensible on these self-reported figures for coding and agentic work. It is not a uniform result, and it comes from a table published by Beam's maker.
The "US has no open-weight option" framing is undercut by the video's own screenshot
The clip's argument for why Beam matters is procurement, not leaderboards: "a lot of my customers cannot get approvals internally to be able to fine tune and serve models that are created in China." That constraint is real and is the strongest point in the video.
The premise attached to it is weaker. Beam is not the first credible US open-weight option to arrive. The second column visible in the frame the creator chose, "Inkling," is Thinking Machines Lab's 975-billion-parameter open-weights MoE with 41 billion active per token, released by a US lab founded by Mira Murati. Nemotron 3 Ultra, the model the video names in its own closing line, is a US open-weight release from NVIDIA that shipped in June 2026. Both appear in Reflection's comparison set. A US-only buyer had options before 2026-10-05, and Beam adds a third rather than opening a closed category.
The China-is-ahead half of the framing holds up better. GLM 5.3 shipped on 2026-08-14 from Z.ai, formerly Zhipu AI. Kimi K3 shipped on 2026-07-16 from Moonshot AI. DeepSeek V4.1 shipped in September 2026. All three predate Beam's announcement, and all three lead it on most published coding and reasoning rows. On causal order, Beam is the follower, which makes "take down" a forecast rather than a measured result.
Reflection AI is not the Reflection 70B of 2024
The name invites a collision worth clearing up, because this story is about self-reported benchmarks. Reflection AI was founded in 2024 by Misha Laskin and Ioannis Antonoglou, both former Google DeepMind researchers. The company emerged from stealth in March 2025 with $130 million, raised $2 billion at an $8 billion valuation in October 2025, and closed a $2.5 billion Series C at a $25 billion pre-money valuation in April 2026, followed by a compute agreement with SpaceX reported at up to $6.3 billion in June 2026. Laskin discussed the Beam launch on CNBC on 2026-10-06.
Reflection 70B was a different thing entirely: a Llama fine-tune announced on 2024-09-05 by Matt Shumer of OthersideAI, promoted as the top open-source model, then found by independent evaluators to be irreproducible, with Artificial Analysis measuring the downloadable Hugging Face weights as roughly on par with Llama 3 and worse than the Llama 3.1 it was supposedly built from. The two have no connection. The relevance here is narrow and specific: the 2024 episode is the reason unreplicated launch-day benchmark tables deserve a label, and the reason Beam's numbers should be treated as provisional until the Apache 2.0 weights land and someone outside Reflection runs them.
One more absence is worth naming. Reflection's comparison set contains only open-weight models. No closed frontier model appears in any of the four tables, so the video's "it's up there" claim says nothing about how Beam compares to the current Anthropic lineup of Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5, and Claude Haiku 4.5, or to anything from OpenAI or Google. The table answers a narrower question than the title implies.
Key Takeaways
- Verified: Beam is Reflection AI's first open-weight model, a text-only sparse MoE with 501B total parameters and 23B active, a 1M-token context, pretrained on 23.8T tokens, announced 2026-10-05. The on-screen scores of 44.4, 77.2, 65.5, 80.1, 34.6, 78.0, and 80.9 all match Reflection's published table exactly.
- Correction: "Today Reflection released Beam" is wrong. Reflection announced Beam and said it will publish weights, technical report, model card, and developer artifacts "later this month" under Apache 2.0. On 2026-10-06 the model is early-access waitlist only.
- Correction: The title's "take down DeepSeek, Kimi, and GLM" is contradicted by Reflection's own table. Beam trails GLM 5.3, Kimi K3, and DeepSeek V4.1 on nearly every agentic coding and reasoning row, including Terminal Bench v2.1 at 80.1 against 88.2, 88.3, and 90.6 respectively.
- Partial correction: "It's up there with GLM and Kimi and DeepSeek" holds only against GLM 5.2, the prior generation. Reflection's own wording is "comparable to GLM-5.2 while using 3-4x less inference compute," not parity with the current frontier.
- Partial correction: "More intelligent than Nvidia's Nemotron Ultra" is mostly supported on coding and agentic rows but is not a sweep. The model is Nemotron 3 Ultra, and it leads Beam on IFBench at 81.7 against 79.7 and ties on AA-LCR at 79.3.
- Unverified: "Some of the smartest engineers in my network joined them at the end of 2025." No public record checked confirms or refutes a specific hiring claim about the creator's personal network.
- Context: Every figure in the table is self-reported by Reflection and unreplicated as of 2026-10-06. The 3-4x efficiency claim excludes prompt prefill, context-dependent attention, and serving overhead by Reflection's own methodology note, so it is an active-parameter estimate rather than a measured serving cost.
- Context: The US open-weight category was not empty before Beam. Thinking Machines Lab's Inkling, the second column visible in the video's own screenshot, and NVIDIA's Nemotron 3 Ultra, named in the video's closing line, are both US open-weight releases that shipped earlier in 2026.
- Context: Reflection AI, founded by Misha Laskin and Ioannis Antonoglou and valued at $25 billion pre-money in April 2026, is unrelated to the 2024 Reflection 70B episode involving Matt Shumer and OthersideAI.
Resources
- Introducing Beam: Reflection's 501B open-weight model. Reflection's own announcement, the source of the headline screenshot at 0:03 and the benchmark table at 0:40, including the full eight-column comparison the video crops and the methodology note on the compute-efficiency estimate.
- Reflection on X: Introducing Beam. Reflection's launch post stating 501B total and 23B active parameters and "Full weights release this month," which is the primary evidence that weights were not out on announcement day.
- Reflection CEO Misha Laskin on launch of new AI model Beam. First-party interview on 2026-10-06 confirming company leadership and launch framing.
- Inkling: Our Open-Weights Model. Thinking Machines Lab's announcement, establishing that the "Inkling" column in the video's screenshot is a US open-weights model, 975B total and 41B active.
- NVIDIA Nemotron 3 Ultra coverage. Establishes the correct model name, the 550B total and 55B active parameter counts, and the 2026-06-04 Computex announcement that puts a US open-weight release four months ahead of Beam.
- Reflection AI, Wikipedia. Founding by Misha Laskin and Ioannis Antonoglou in 2024 and the funding history used for the valuation figures.
- New open source AI leader Reflection 70B's performance questioned, accused of 'fraud'. Contemporaneous reporting on the separate 2024 Reflection 70B episode, included to mark the name collision and to ground why unreplicated launch benchmarks warrant a label.
Published October 5, 2026. Writeup generated from a favorited TikTok.