The 3x Token Claim Holds Up: Anthropic's Own Sweep Shows Claude Fable 5.1 Going From 73k to 222k Median Tokens Between Low and Max Effort, but Hardware Gained 41 Points While Security Gained 23
Watch on TikTok
The file is a 67.93-second MP4 at 720x1280, 30 fps measured with ffprobe (the fps field in the yt-dlp metadata is null), HEVC video and AAC stereo audio at 44.1 kHz, overall container bitrate 189,383 bps against a reported tbr of 189, audio stream 48,115 bps, 1,608,185 bytes on disk (1.53 MiB), SDR, carrying the format id bytevc1_720p_189382-1 - 720x1280. It posted 2026-10-05T22:43:26 UTC from the handle @forgedlabsai, whose channel nickname is Forged Labs, over an audio track listed as "original sound" and credited to the artist "Forged Labs." At the 2026-10-07 capture, two calendar days and roughly 25 hours after posting, the video showed 369 views, 13 likes, 0 comments, 3 reposts, and 6 saves. I read all 34 frames in tokdoc_frames/ and all 130 words of transcript.txt. Every frame carries the same composition: a full-bleed dark-mode screenshot of a line chart with a webcam picture-in-picture pinned to the lower right. The chart header reads "Terminal-Bench 3.0" with the subhead "Pass rate against tokens spent, by effort setting" and a small "PNG" label at the top right. The y-axis is labeled "Share of attempts that passed" with gridlines at 20%, 30%, 40%, 50%, 60%, and 70%; the x-axis label reads "Median tokens pe" before the PiP cuts it off, with visible ticks "50k" and "100". Four series run bottom-left to top-right and are labeled at their right-hand endpoints "Opus 5.5 / max" (solid white, highest), "Fable 5.1 / max" (salmon), "Opus 5 / max" (dashed grey), and "Fable 5 / max" (dotted grey). Individual points carry the effort labels "low", "medium", "high", "xhigh", and "max". A two-line annotation sits near the 60% gridline reading "same score, / half the tokens", positioned beside the Opus 5.5 "high" point and level with the Fable 5.1 "max" point. Frames 001 and 002 alone add a white rounded card over the chart's upper third with purple text reading "Don't pay AI to / Overthink Everything"; the card is gone by frame 003 and never returns. Burned-in captions advance through the narration with single keywords highlighted in yellow: "big money" (frames 001-002), "anthropic" (005-007), "three times" (008-010), and "useful" (029). The speaker is a bald man with short grey stubble, a grey crewneck tee, a lavalier capsule visible at his collar in frames 029-030, an in-ear bud, and a forearm tattoo visible in frames 008 and 033, lit by a soft frontal key against a plain interior. The caption text across the full set reads, in order: "most companies are about to waste big money" (001-002), "by asking AI to think deeply about everything" (003-004), "an anthropic engineer just tested different AI effort levels" (005-007), "and found that higher effort used roughly three times as many tokens" (008-010), "it substantially improved work with the hidden edge cases" (011-012), "especially security" (013), "but delivered much smaller gains on some routine tasks" (014-015), "the business lesson is simple" (016), "match AI effort to the cost of being wrong" (017-018), "use fast inexpensive settings for brainstorming" (019-021), "drafts and early iterations" (022), "and increase the effort for financial analysis" (023-024), "security reviews customer commitments" (025), "compliance and final verification" (026-027), "my verdict this is useful now" (028-029), "don't give every task the same AI budget" (029-030), "spend more reasoning where errors create real consequences" (031-032), and "not where an employee can quickly review the result" (033-034).
The chart on screen is Anthropic's, the author is an Anthropic engineer, and the three-times-tokens figure is accurate
The narration says, "An anthropic engineer just tested different AI effort levels and found that higher effort used roughly three times as many tokens in one benchmark." Both halves check out.
The chart in every frame is reproduced from Spending your effort, published on claude.dev on September 25, 2026 and bylined Thariq Shihipar. The site's own header describes itself as a "Blog / technical writing for people building with Claude" and the tagline reads "Sharing tips, tricks, and POVs from Anthropic's developers," which I confirmed by fetching claude.dev on 2026-10-07. Shihipar's speaker bio on ai.engineer, fetched the same day, states he is a "member of technical staff at Anthropic working on Claude Code." The video's "anthropic engineer" description is correct, and the chart's "Terminal-Bench 3.0" header matches the article's figure exactly, including the "same score, half the tokens" annotation visible in frames 003 through 034.
The token multiple comes from the article's run counts. It reports Claude Fable 5.1 at low effort as "370 attempts, median 73k tokens each" and at max effort as "370 attempts, median 222k tokens each." That is 3.04 times the median token spend. "Roughly three times" is right to two significant figures. The 370 attempts also reconcile cleanly with the benchmark's size: Terminal-Bench 3.0 shipped in July 2026 with 74 tasks, per the Snorkel AI leaderboard writeup, and 74 tasks at five attempts each is 370.
What the video drops is which model the figure describes. The 73k-to-222k sweep is Claude Fable 5.1, not any Opus model, and it is low effort against max effort rather than a general statement about "higher effort." The chart on screen plots four separate models with four separate curves, and the narration never names any of them.
Hardware, not security, posted the largest gain in the article's own domain table
Frame 013 pairs the caption "especially security" with the chart still on screen. The article does break out pass rates by domain at low effort versus max effort, and the numbers are:
| Domain | Low | Max | Change |
|---|---|---|---|
| Security | 64% | 87% | +23 points |
| Hardware | 34% | 75% | +41 points |
| ML | 54% | 73% | +19 points |
| Science | 41% | 61% | +20 points |
| Software | 43% | 56% | +13 points |
| Media | 18% | 30% | +12 points |
| Operations | 12% | 22% | +10 points |
Security improved, so the direction of the video's claim is right. The superlative is wrong. Hardware gained 41 percentage points, nearly double security's 23, and hardware more than doubled its pass rate while security rose by about a third of its starting value. If the video wanted one headline domain, the data points at hardware.
Where the security framing does hold is in the article's worked example. It names html-js-filter, "a Terminal-Bench 3.0 task that asks for an HTML sanitizer that strips every way of smuggling JavaScript into a page," and reports that "Fable 5.1 went from 1/5 at low to 5/5 at xhigh." That is a security-flavored task with a clean effort response, and it is likely what the video is compressing. One task is not the same thing as a domain ranking.
"Much smaller gains on some routine tasks" is a fair read of the absolute numbers and a misleading read of the relative ones
The transcript says higher effort "delivered much smaller gains on some routine tasks." Operations and media are the two smallest absolute movers in the table, at +10 and +12 points. On that measure the claim stands.
Two things complicate it. First, the article never calls operations or media "routine." They are the two domains where both models score worst in absolute terms, which is a different problem from the task being easy. Operations tops out at 22% even at max effort. Second, in relative terms operations rose from 12% to 22%, which is an 83% increase in pass rate, and media rose from 18% to 30%, a 67% increase. Those are larger relative movements than security's 64% to 87%. The video picks the framing that supports its thesis without saying which framing it picked.
The chart's loudest claim is one the narration never makes
The annotation sitting at the 60% gridline in frames 003 through 034 reads "same score, half the tokens." On the chart it is placed beside the Opus 5.5 "high" point and level with the Fable 5.1 "max" point, which sits at roughly twice the x-position. The article frames this as the generational improvement: "Fable 5.1 and Opus 5.5's effort curves are our best yet: at each level, there is an uptick in benchmark scores and tokens consumed."
That is a claim about picking a better model, and it is independent of the video's thesis about picking a lower effort level. Vendor pricing makes the gap wider than the token count alone. Per Anthropic's pricing page, fetched 2026-10-07, Claude Opus 5.5 is $4/MTok input and $20/MTok output, while Claude Fable 5.1 is $10/MTok input and $50/MTok output. Opus 5.5 costs 40% of Fable 5.1 per token on both sides. Anthropic's Opus 5.5 launch page, fetched the same day, states that "Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost." A viewer who acts only on the video's advice tunes the effort dial and leaves the model choice untouched, which is the larger of the two levers on this chart.
One caveat on reading any token count across model generations: the same pricing page notes that "Claude 4.7 and later models and Claude Mythos Preview use a newer tokenizer" that "produces approximately 30% more tokens for the same text." Every model on this chart is post-4.7, so the comparison between them is internally consistent.
The source tested 74 terminal coding tasks; the video extends it to compliance and customer commitments
Frames 023 through 027 carry the operational advice: "increase the effort for financial analysis," "security reviews customer commitments," "compliance and final verification." The posted description repeats the list, naming "security, financial decisions, compliance, and final verification."
None of that is in the source. The article's domains are hardware, science, machine learning, software engineering, security, operations, and media, all measured by whether an agent completes a terminal task. Terminal-Bench is described by its maintainers as "a benchmark to measure and evolve with the frontier of agent work," hosted by Stanford, Harbor, and the Laude Institute, and the Snorkel writeup describes it as evaluating whether agents can operate real terminals, "writing and debugging code, configuring environments, and recovering from failure, all without a GUI." There is no financial analysis task, no compliance task, and no customer-commitment task in it.
The article's own guidance is written for coding and stays there. Low is "for when I want quick responses that are in the loop, e.g. brainstorming, sketching, easy changes." Medium is "for most of my regular software engineering work, e.g. new feature implementation." High is "for work where verification is important or there are edge cases, e.g. fixing a bug in a brownfield codebase." Max is "When I want Claude to operate fully autonomously to solve difficult problems, e.g. end to end building and verification of an app."
The video's extension to business workflows is a reasonable analogy. It is presented as a finding. The opening line, "Most companies are about to waste big money by asking AI to think deeply about everything," attaches no spend figure, no company count, and no source, and nothing in the article supports it.
This post is not selling anything I could find
The description names no product, carries no link, and ends with hashtags only: #ArtificialIntelligence #AICosts #BusinessStrategy #AILeadership #Productivity. The transcript contains no call to action. The handle @forgedlabsai is unrelated to the Minecraft creator "Forge Labs" (@forgelaboratories) that dominates search results for that name. I fetched the uploader's profile page on 2026-10-07 and it returned only TikTok's generic "Make Your Day" shell with no bio, link, or follower count, so I cannot confirm what, if anything, the account promotes elsewhere. On the evidence in this video, it is commentary rather than an ad.
Key Takeaways
- Verified: The "roughly three times as many tokens" figure is accurate. Anthropic's article reports Claude Fable 5.1 on Terminal-Bench 3.0 at a median 73k tokens per attempt at low effort and 222k at max effort across 370 attempts each, a ratio of 3.04.
- Verified: The "anthropic engineer" attribution is correct. Thariq Shihipar is listed as a member of technical staff at Anthropic working on Claude Code, and claude.dev describes itself as carrying POVs from Anthropic's developers.
- Verified: Higher effort did substantially improve the security domain, from 64% to 87% pass rate, and the article's
html-js-filterHTML sanitizer task went from 1/5 at low effort to 5/5 at xhigh. - Correction: Security was not the standout domain. Hardware gained 41 percentage points (34% to 75%) against security's 23. The video's "especially security" picks the second-largest gain.
- Correction: The video presents the 3x figure as a property of "higher effort" in general. It is one model (Claude Fable 5.1) at the two extremes of a five-level scale (low versus max) on one benchmark. The chart on screen plots four different models with four different curves and the narration names none of them.
- Partial correction: "Much smaller gains on some routine tasks" is true in percentage points for operations (+10) and media (+12), and false in relative terms, where operations rose 83% and media 67% from their starting pass rates. The article does not label either domain routine.
- Partial correction: The advice list in frames 023-027 (financial analysis, customer commitments, compliance) describes work the source never tested. Terminal-Bench 3.0 is 74 terminal coding tasks. The extrapolation may be sound; it is not a finding from this data.
- Unverified: "Most companies are about to waste big money." No spend figure, no company count, and no source appear in the video or in the article it draws from.
- Context: The chart's own annotation, "same score, half the tokens," is about choosing Opus 5.5 over Fable 5.1, not about choosing a lower effort level. Opus 5.5 is also priced at 40% of Fable 5.1 per token on both input ($4 versus $10 per MTok) and output ($20 versus $50 per MTok).
- Context: Anthropic's pricing page notes that Claude 4.7 and later use a tokenizer producing about 30% more tokens for the same text. Every model on this chart is post-4.7, so the token comparisons between them are consistent.
- Context: The video names no product and carries no link. The uploader's profile page would not load any bio or link for me on 2026-10-07, so I cannot rule out promotion elsewhere on the account.
Resources
- Spending your effort, Thariq Shihipar, claude.dev, September 25, 2026: the source of the on-screen chart, the 73k/222k token medians, the 370-attempt counts, the per-domain pass-rate table, the
html-js-filterexample, and the low/medium/high/max guidance. - claude.dev: establishes the blog's own description as "technical writing for people building with Claude" sharing "POVs from Anthropic's developers."
- Thariq Shihipar speaker bio, ai.engineer: confirms he is a member of technical staff at Anthropic working on Claude Code.
- Anthropic pricing documentation: the live per-million-token prices for Claude Opus 5.5 ($4/$20), Claude Opus 5 ($5/$25), Claude Fable 5.1 ($10/$50), and Claude Fable 5 ($10/$50), plus the note on the newer tokenizer producing about 30% more tokens.
- Introducing Claude Opus 5.5, Anthropic: confirms the model name, the low/medium/high/xhigh effort levels with medium as default, and the statement that Opus 5.5 at default effort beats Opus 5 at max effort for about a fifth of the cost.
- Terminal-Bench 4.0 leaderboard writeup, Snorkel AI: establishes that Terminal-Bench 3.0 shipped in July 2026 with 74 tasks and that the benchmark measures agents operating real terminals without a GUI.
- Terminal-Bench: the benchmark's own site, confirming it is hosted by Stanford, Harbor, and the Laude Institute.
- @forgedlabsai on TikTok: the uploader's profile; it loaded but returned no bio, link, or follower data, which is why the promotion question stays open.
Published October 5, 2026. Writeup generated from a favorited TikTok.