<- all tokdocs

Karpathy's four-format ladder checks out against the original post, but the spec is ASD-STE100 Issue 9 and the body that owns it warns AI text can look compliant without being compliant

Watch on TikTok

View on TikTok ->

A 102-second vertical clip (102.119909 s measured by ffprobe, 1:42 as reported in metadata), 720x1280 at 30 fps, HEVC Main profile video at 842,183 bps, AAC HE-AACv2 audio at 44,100 Hz stereo and 32,062 bps, 880,798 bps overall, 11,243,384 bytes (10.72 MiB) in an MP4 container, SDR, posted 2026-10-02 at 14:01:49 UTC by @agenticengineering under the channel nickname "Agentic Engineering" with the audio track labeled "original sound" and credited to the same account. Captured 2026-10-07, five days after posting, the clip showed 510 views, 24 likes, 5 comments, 10 reposts and 21 saves. I read all 51 frames in tokdoc_frames/ (sampled at 2-second intervals, covering 0 s to 100 s) and all 267 words of transcript.txt. The frames show one continuous handheld selfie take with no cuts, no screen recordings, no diagrams and no B-roll, which matters because a video arguing for visual explanation contains zero visuals beyond the speaker's face. Frames 1 through 7 carry a single on-screen caption, white bold text on a rounded light-blue box pinned to the top of the frame, reading verbatim "Stop asking AI to explain things with just more text". That caption is gone by frame 8 and never returns; the remaining 44 frames carry no text at all. The speaker is a man with short dark hair greying at the temples and a close-cropped grey-flecked beard, wearing an open-collar ribbed grey-blue short-sleeve polo, walking downhill while holding the camera at arm's length slightly below eye level. The setting is a pale limestone street under bright midday sun: a white-and-glass domed rooftop structure sits behind his left shoulder in frames 1 through 20, magenta and purple bougainvillea spills over a low stone retaining wall on the left, black wrought-iron street lamps and a broad flight of pale stone steps occupy the right, and arched windows with iron balconies run along a sandstone building behind him. Cumulus clouds build across the frame as he walks. By frame 23 the dome has slipped out of shot, by frames 31 to 36 a green hedge and tree line enter from the right, and from frame 37 to frame 51 a dense dark-green conifer fills the right half of the frame while the cloud cover thickens and the paving changes to flagstone. His free hand enters the bottom of the frame to gesture in frames 13, 30 and 43. The transcript's final line, "Let me know what you think and if you have tried it out," appears twice, and transcript.srt cue 18 is timestamped 00:01:41,740 to 00:02:11,720, which runs 29.6 seconds past the end of a 102-second file. That is a Whisper repetition artifact, not a second utterance, and it inflates the raw word count.

The source is a post by Andrej Karpathy, not "Andrew Karpathy," and it went up 13 hours and 24 minutes before this video

The transcript opens: "Andrew Karpathy shared a really useful progression for getting better and better explanations out of large language models." The name is wrong. The post is by Andrej Karpathy, at x.com/karpathy/status/2105819303471976479. I retrieved its full text through the fxtwitter API mirror, which returns the canonical permalink, the author handle @karpathy, and a creation timestamp of Fri Oct 02 00:37:00 +0000 2026. This TikTok was posted at 2026-10-02 14:01:49 UTC, 13 hours and 24 minutes later. US coverage dated the post to October 1 because 00:37 UTC on October 2 falls on the evening of October 1 in US Pacific time.

The four rungs in the video match the four rungs in the post, in the same order. Karpathy's exact text runs: "Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100 ... But even better: Diagrams / images. Instead of writing, ask your LLM to create a diagram ... But even better: Web pages. Ask for output 'in HTML' to get a beautiful, interactive webpage ... But even better: Explainer videos."

The video drops three specifics. Karpathy named the concrete prompt for the fourth rung: "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration." He flagged the dependency plainly, writing "(you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute)." And he closed with a framing the video never mentions, that you can now ask for "large, custom, discardable software artifacts ... that would have never made sense to create before." The video's rendering, "let the model build the whole thing," removes the API key requirement and the discardability point, which are the two things that determine whether the fourth rung is practical for a given viewer. For reference, 3Blue1Brown is Grant Sanderson's mathematics animation channel, and Manim is the Python animation library he originally wrote, now maintained by the Manim Community under the MIT license.

The specification is ASD-STE100, Issue 9, dated January 2025, and it is free

The transcript renders the name as "ASD-ST100" in both places it appears. The correct designation is ASD-STE100, with the E for English. The video's description of it holds up against the primary source. It says the spec is "a controlled English specification that was originally developed for aerospace maintenance documentation" that "limits vocabulary and sentence structure."

The official ASD-STE100 about page, run by the ASD Simplified Technical English Maintenance Group, states that STE "is a controlled natural language developed in the late 1970s (originally as AECMA Simplified English) to help the users of English-language maintenance documentation understand what they read," that it "was initially applicable to commercial aviation," and that it "was first released in 1986 as AECMA Document, PSC-85-16598." The official FAQ dates the project start to 1979, AECMA's decision to build its own controlled language to 1981, and the formal start of the AECMA Simplified English Working Group to June 30, 1983, at the Fokker plant near Schiphol Airport.

Three details the video does not give, each of which changes what a viewer should do next. First, the current version is Issue 9, dated January 2025, and with Issue 9 the document stopped being a specification and became an international standard. Legacy pages still live on the same domain at /about.html and /faq.html and still say Issue 8, April 2021, with "the next issue 9 is scheduled in 2024," so anyone checking a stale link will read an outdated issue number. Second, the controlled dictionary holds "approximately 900 approved words," per the official FAQ. Third, the standard costs nothing: the FAQ states it "is available to everyone free of charge" in PDF, and the downloads page carries the request form. The video tells you to invoke the spec by name and never mentions that you can read the thing itself in an afternoon.

ASD's own maintenance group says STE was not built for this use, and that AI output can read as compliant without being compliant

This is the finding the video misses entirely, and it comes from the body that owns the standard.

On the use question, the official FAQ answers "Who needs to write in STE?" with: "STE was developed to make maintenance documentation easier to read, so writers of such documentation use it in procedural and descriptive texts. It is not intended for general-purpose writing, such as international correspondence. However, many of its principles (for example, short sentences, one topic per sentence, and the use of the active voice) can be effectively adopted and applied in other writing contexts." The same page answers "Can STE be used to teach English?" with "No." Asking a model to explain a concept in STE falls outside the standard's stated scope. The principles transfer; the standard was not written for the job. Karpathy's own hedge about asking for "80% of the way to ASD-STE100 because the spec is quite stringent" is consistent with that, and the video reproduces the hedge accurately: "Karpathy says you can even ask for something like 80% of the way to the [ASD-STE100] if the full specification feels too rigid."

On the AI question, the STEMG and its Artificial Intelligence Task Team published a white paper on ASD-STE100 and AI dated June 2026, four months before this video. It lists "Automated checks for STE compliance, but currently with varying levels of accuracy" as a benefit with a caveat attached, and lists "Variable reliability and factual accuracy of AI-generated content" and "Potential compromise of terminology control" under challenges. The download page summarising it is blunter: "AI-generated text can appear clear, authoritative, and consistent with STE, even when it does not correctly apply the rules and vocabulary of the standard. Plausibility must not be confused with verified compliance." The white paper's conclusion sets the condition: "AI should assist, not replace, human authors."

That is a direct limit on the video's first rung. The transcript says the spec "limits vocabulary and sentence structure to make technical writing clearer and less ambiguous," which is true of the standard. It is not automatically true of a model asked to imitate the standard. The standard's own custodians say the imitation can pass a reading test while failing a compliance test.

The "a good visual beats five paragraphs" claim has real research behind it, but that research measured students learning how machines work

The transcript's argument for the second rung is: "For a lot of technical concepts, a good visual is simply easier to understand than another five paragraphs." The video cites no evidence, and Karpathy's post cites none either. He wrote only that diagrams "can be a lot easier to process, parse, and understand."

There is supporting research, and the numbers are specific. Richard Mayer, in his 2023 chapter "Research-Based Principles for Designing Multimedia Instruction" for the American Psychological Association Division 2 volume In Their Own Words, writes: "In a series of 13 experimental comparisons my colleagues and I have found that students perform much better on a transfer test when they learn from words and graphics than from words alone (e.g., narration and animation versus narration alone, or text and illustrations versus text alone), yielding a median effect size of d = 1.35." His worked example is a narrated animation of a tire pump, tested with troubleshooting questions, from Mayer and Anderson (1991).

Two things keep that from being a clean endorsement of the video's claim. The materials were short instructional lessons on physical and scientific systems (tire pumps, brakes, generators, lightning) delivered to students, with transfer tests as the outcome. The population and the task are not an adult reviewing a language model's answer to a question they posed. And the broader literature is less dramatic than d = 1.35. The 2025 meta-analysis by Jennifer G. Cromley and Runzhi Chen, "A meta-analysis of Richard Mayer's multimedia learning research: Searching for boundary conditions of design principles across multiple media types", Educational Research Review vol. 49, DOI 10.1016/j.edurev.2025.100730, synthesised 92 peer-reviewed articles containing 181 studies and 591 effects and reported an "overall effect was g = 0.37," moderated by every moderator tested, including a small decline in effect size per year. Within that, text combined with diagrams held up well, while virtual reality produced no significant effects. So the second rung of the ladder is the one with the strongest empirical backing, and the honest version of the claim is a medium effect across a varied literature rather than a universal "simply easier."

The post this video compresses had 7,516,814 views; the video had 510

At the time I checked on 2026-10-07, the Karpathy post carried 7,516,814 views, 54,202 likes, 6,365 reposts and 1,559 replies. This TikTok, five days after posting, carried 510 views, 24 likes, 10 reposts, 21 saves and 5 comments.

The compression ratio matters more than the reach gap. A roughly 320-word post became a 267-word spoken summary that kept the four rungs and the 80 percent hedge intact, misread the author's first name, lost the spec's E, and dropped the API key dependency that gates the fourth rung. Nothing in the video is invented. The errors are all errors of transmission, which is the failure mode you get when a claim travels one hop from its source without the source attached. The video never shows or links the post. A viewer who wants to check it has to search for a name that is spelled wrong and a standard that is named wrong.

Key Takeaways

  • Correction: The transcript says "Andrew Karpathy." The author is Andrej Karpathy, and the post is x.com/karpathy/status/2105819303471976479, created Fri Oct 02 00:37:00 +0000 2026.
  • Correction: The transcript renders the standard as "ASD-ST100" twice. The correct designation is ASD-STE100, per the official site.
  • Correction: The current version is Issue 9, dated January 2025, and it is now an international standard rather than a specification. Legacy pages on the same domain still say Issue 8, April 2021.
  • Verified: The four-rung ladder in the video (controlled English, then diagram, then HTML page, then explainer video) matches the order and content of Karpathy's post, including the "80% of the way to ASD-STE100" hedge and his stated reason, "because the spec is quite stringent."
  • Verified: STE was developed in the late 1970s as AECMA Simplified English for commercial aviation maintenance documentation and first released in 1986 as AECMA Document PSC-85-16598. The video's "originally developed for aerospace maintenance documentation" is accurate.
  • Verified: The standard is free. The official FAQ states it "is available to everyone free of charge" in PDF, requestable from the downloads page. The controlled dictionary holds approximately 900 approved words.
  • Context: ASD's own FAQ states STE "is not intended for general-purpose writing," which puts the video's use case outside the standard's stated scope. The principles transfer; the standard was written for procedural and descriptive maintenance text.
  • Context: The STEMG white paper of June 2026 warns that "AI-generated text can appear clear, authoritative, and consistent with STE, even when it does not correctly apply the rules and vocabulary of the standard," and that "AI should assist, not replace, human authors." Neither the video nor the post mentions this.
  • Partial correction: "A good visual is simply easier to understand than another five paragraphs" overstates what the research supports. Mayer reports a median d = 1.35 across 13 comparisons, measured on students taking transfer tests about physical systems. The 2025 Cromley and Chen meta-analysis of 92 articles, 181 studies and 591 effects reports an overall g = 0.37 with significant moderation.
  • Unverified: The claim that "the model can increasingly build the explanation around how you learn best" has no cited evidence in the video, none in Karpathy's post, and no primary source I could locate. Treat it as an assertion.
  • Context: The video omits Karpathy's concrete fourth-rung prompt, "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration," and his note that you need an API key or a local-compute alternative. ElevenLabs does require an API key for its text-to-speech API.
  • Context: The transcript's 267-word count includes a duplicated closing line; transcript.srt cue 18 runs to 00:02:11,720 in a 102-second file, which is a Whisper repetition artifact.

Resources

  • api.fxtwitter.com mirror of Karpathy's post returns the full verbatim text, the canonical permalink https://x.com/karpathy/status/2105819303471976479, author @karpathy, timestamp Fri Oct 02 00:37:00 +0000 2026, and engagement counts of 54,202 likes, 6,365 reposts, 1,559 replies and 7,516,814 views; the x.com page itself returns HTTP 402 to automated fetches.
  • About ASD-STE100, official STEMG site establishes the late-1970s origin, the 1986 release as AECMA Document PSC-85-16598, and that the current version is Issue 9, January 2025.
  • ASD-STE100 official FAQ establishes the ~900-word approved dictionary, the free-of-charge distribution, the 1979 project start and 1983 working group founding, and the statement that STE "is not intended for general-purpose writing."
  • ASD-STE100 downloads page establishes that Issue 9 (January 2025) is the copy you can request for free, and carries the summary line that plausibility must not be confused with verified compliance.
  • STEMG white paper, ASD-STE100 and Artificial Intelligence dated June 2026, establishes the standard owner's position that AI must assist rather than replace human authors, and lists variable factual accuracy and compromised terminology control as risks.
  • Mayer, "Research-Based Principles for Designing Multimedia Instruction" (2023) establishes the multimedia principle, the 13 experimental comparisons and the median effect size of d = 1.35 on transfer tests.
  • Cromley and Chen (2025) meta-analysis record, NSF PAR establishes the full citation in Educational Research Review vol. 49, DOI 10.1016/j.edurev.2025.100730, the 92 articles / 181 studies / 591 effects corpus, and the overall g = 0.37.
  • Manim Community establishes that Manim is an MIT-licensed Python animation library originally written by Grant Sanderson and now community-maintained.
  • 3Blue1Brown establishes that the channel referenced by Karpathy's "3b1b" shorthand is Grant Sanderson's mathematics animation channel.
  • ElevenLabs API reference establishes that the text-to-speech API authenticates with an API key, confirming the dependency Karpathy named and the video omitted.

Published October 2, 2026. Writeup generated from a favorited TikTok.