<- all tokdocs

Papermorph Ships 88 Real Chapters, but Its Voice Is Microsoft Edge TTS and "No Image Models" Is a Limitation the Author Plans to Remove

Watch on TikTok

View on TikTok ->

The artifact is a 40.67-second clip (0:40), 720x1280 portrait at 30 fps, HEVC Main video at roughly 106 kbps paired with HE-AACv2 audio at 64 kbps and 44.1 kHz stereo, 902,896 bytes on disk (881.7 KB), overall bitrate about 178 kbps. It was posted 2026-10-05 at 07:50:01 UTC by the handle @github.signals under the channel nickname "Github Signals" (numeric uploader_id 7614639676569895957). At capture on 2026-10-06 it showed 24,000 views, 1,079 likes, 5 comments and 303 reposts. The audio track is listed as "original sound" by Github Signals, meaning a synthetic voiceover rather than a licensed music bed. I read 12 of the 20 extracted frames (001, 003, 005, 007, 009, 010, 011, 012, 014, 016, 018, 020) against a 106-word transcript. The footage is a screen recording of a desktop Chrome window inside a portrait frame. The address bar reads "github.com/dozentwelve/papermorph" and the tab title reads "GitHub - DozenTwelve/Papermorph: An AI skill that turns books". The README renders with a blue-and-black "Papermorph" wordmark and the tagline "A Skill that turns PDFs into animated interactive web books." Below it: "You've seen Opus 5.5 one-shot videos. This Skill takes it further: books you can explore, listen to, and interact with." Two links read "Explore the live bookshelf →" and "Use the Skill". A dark-green embedded demo plays, captioned "Watch the full 75-second demo with narration". A monospace code block spells the pipeline: "PDF -> Book plan -> Storyboards -> Narration -> Animation & quizzes". Three bolded README lines are legible across frames 001 through 011: "Today: Opus 5.5 only. No image models, multilingual support, or BGM yet.", "Planned: Image models and storyboarding for interactive picture books and humanities documentaries.", and "Milestones: Expand the bookshelf—from STEM textbooks to picture books and social science titles—and release new Skills." Under "Get started" the install command reads "npx skills add DozenTwelve/Papermorph --skill papermorph --agent c..." followed by a prompt block beginning "/papermorph Turn /path/to/book.pdf into an animated interactive web book. Target readers: [your audience]. Start with one English chapter for review." From frame 009 the recording switches to the full-screen preview file "papermorph-preview.mp4", showing a wooden bookshelf with spines labeled "Elementary Algebra" (subtitle "68 narrated chapters", cover equation a²-b²=(a-b)(a+b)), "Elementary Mathematics", and "Elementary Physics" marked "Physics · Coming soon", with a lower shelf of "Little Bear in the Woods", "Where Are the Stars?" and "Why Is It So Fun?". Later frames show a chapter cover titled "AN ANIMATED WORKBOOK / Elementary Algebra", a contents page listing "UNIT 1 Arithmetic properties" with numbered entries "1 Types of numbers", "2 Algebraic properties", "3 Order of operations" and "UNIT 2 The number system" with "4", "7 Multiplying and dividing fractions", "8 Adding and subtracting fractions", "10 Multiplying and dividing decimals", and a lesson stage drawing "3 + 5 = 8" and "5 + 3 = 8" as two colored arcs over a 0-to-9 number line with a player scrubber reading "2 / 22 Order in addition". Frame 020 is the outro card: "Follow for more open-source projects" over "@GithubSignals".

The project is real, shipping, and larger than the clip suggests

Papermorph is not a paper, a mock or a teaser. The canonical repository is DozenTwelve/Papermorph, MIT licensed, primary language HTML, created 2026-10-03T01:38:22Z and last pushed 2026-10-05T12:33:10Z. At my check on 2026-10-06 it carried 385 stars and 42 forks with zero open issues. The repository tree holds 1,621 files, of which 1,406 are MP3 narration clips. The live site at papermorph.diamonddoge.org lists two finished titles with the exact labels "Elementary Algebra ↗ 68 narrated chapters" and "Elementary Mathematics ↗ 20 narrated chapters". That is 88 completed chapters of generated HTML and audio, and the repo tree confirms the file count: 68 site/elementary-algebra/chNN/index.html files and 20 under site/math-notebook/.

Two days separate the repository's creation from the video's post. The clip was describing a project that was 48 hours old and already past 300 stars.

The narration does not come from the AI model, it comes from Microsoft Edge

This is the clip's biggest gap. The transcript says Papermorph "turns static PDFs into living, narrated web experiences" and then claims "it does this entirely by generating code and storyboards without needing any image generation models." The framing invites you to conclude that one model produces the whole artifact including the voice.

It does not. The narration is synthesized by Microsoft's Edge text-to-speech service through the edge-tts Python package. The docstring at the top of scripts/tts.py states it plainly: "Generate narration audio and word-mark timings with Edge TTS." The script calls edge_tts.Communicate(text, voice, rate=rate, boundary="WordBoundary"), streams the MP3 to disk, and uses the WordBoundary events to build timings.js so each animation beat can wait for a specific spoken word. SKILL.md lists the hard requirements: "Needs: uv, ffmpeg/ffprobe, Edge TTS network access and Playwright Chromium."

The architecture makes sense once you check what Claude models actually emit. Anthropic's models overview states that "All current models support text and image input, text output, multilingual capabilities, vision, and tool use." Text output only. No current Claude model produces audio, so a project that wants narration has to reach outside the model for it. Edge TTS is the free, no-API-key choice, which is why it is here. The dependency costs nothing in dollars but it does require network access to a Microsoft endpoint, and it is the single component in the pipeline that is neither open source nor under the author's control.

"No image generation models" is a stated limitation, not the headline feature

The transcript calls this "the most surprising part" and presents it as a deliberate engineering win. The README treats it as a gap on the roadmap. The exact line visible on screen in frames 001, 003, 005, 007 and 011 reads:

Today: Opus 5.5 only. No image models, multilingual support, or BGM yet.

And immediately below it:

Planned: Image models and storyboarding for interactive picture books and humanities documentaries.

The word "yet" is doing the work. The author lists missing image models alongside missing multilingual support and missing background music, three things he intends to add. The video inverted a to-do list into a selling point. The underlying technical fact is still true and still interesting: chapter visuals are hand-coded SVG, and SKILL.md instructs "Default to code/SVG artwork; record any other asset choices in the book's conventions." Each chapter is "one 1600×900 SVG stage plus controls". But calling that a decision to avoid image models misreads the source.

The model named on screen checks out, and it is the expensive one

The README says "Opus 5.5 only" and the install instructions say "In Claude Code with Opus 5.5, run:". Checked against the live models overview, Claude Opus 5.5 is current and real, API ID claude-opus-5-5, and the docs recommend it as the starting point: "If you're unsure which model to use, start with Claude Opus 5.5 for most workloads." The rest of the current lineup is Claude Fable 5.1 (claude-fable-5-1), Claude Sonnet 5.5 (claude-sonnet-5-5) and Claude Haiku 4.5 (claude-haiku-4-5), with Claude Opus 5 and Claude Sonnet 5 listed as legacy but still available.

The cost side is worth stating because the video does not. The skill is MIT licensed and free to download, but Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens. The 88 shipped chapters are HTML files in the 22 KB to 27 KB range each, written one at a time through a storyboard-then-implement loop, with SKILL.md recommending "one fresh subagent per chapter sequentially". Nothing about the clip's "you just feed it a book" suggests the token bill behind 88 chapters of that.

The quiz claim is accurate and the mechanism is in the engine

The transcript says you "take quizzes right in your browser," and this one holds up without qualification. The engine reference documents a quiz call used inside a beat:

quiz(position, [{ id: 'c-mean', prompt: ['Find the mean of ', $m('7, 3, 9'), '.'], build: blanks([...]) }, …], done, 'Quick check')

Question builders include choice(options, right, why, wrongWhys[], onRight?), tap and blanks, with grading handled client-side through api.grade(ok, message, {right, total}). Positions include SCREEN for "full-screen practice; label 'Chapter practice N of M'". The shipped site/elementary-algebra/ch01/index.html contains a live instance: a beat named 'q1', 'Quick check' calling quiz(BAND, [{ id: 'c-zero', prompt: ['On the number line, tap the whole number that is ', h('b', '', 'not'), ' a natural number.'], build: tap([0, 1, 2, 3, 4, 5], 0, 'Zero is a whole number, b... It is static HTML with no backend, and progress stays local. The bookshelf's own note confirms it: "Reading progress stays in your browser."

"You just feed it a book" understates the human loop by a wide margin

SKILL.md describes a six-stage pipeline with explicit user checkpoints, not a drop-and-wait. Stage 1, intake, says to "Record readers, tone, primary language, scope and assets/guide. Ask for missing decisions together, with defaults." Stage 3 ends with "Confirm the list and visual approach with the user." Stage 5 is a pilot: "Make chapter 1 through the chapter loop. Deliver it for the user's review and record approved conventions in BOOK.md." Step 5 of the per-chapter loop reserves judgment for a person: "The user reviews overall aesthetics, pacing and teaching effectiveness." The README's own suggested prompt ends with "Start with one English chapter for review."

The prerequisites are also non-trivial for the StudyTok and TeacherTok audience the hashtags target: uv, ffmpeg, Playwright Chromium installed via uv run --with playwright playwright install chromium, a Claude Code install, and an Opus 5.5 budget.

Two shelves in the demo are previews, not products

Frame 009 shows a bookshelf with six spines. Three of them, "Little Bear in the Woods", "Where Are the Stars?" and "Why Is It So Fun?", read as existing children's titles. They are not. The live bookshelf labels those rows "Picture books for children · Coming soon" and "Curious books · Coming soon", and the site's about panel is explicit: "Books on the future shelves are previews, not available lessons." The physics spine carries the same caveat, visible in the frame as "Physics · Coming soon". The clip never claims these are finished, but the shelf imagery reads as a catalog when only two of the six titles exist.

Key Takeaways

  • Verified: The repository is real and substantial. DozenTwelve/Papermorph, MIT licensed, created 2026-10-03, 385 stars and 42 forks at 2026-10-06, 1,621 files including 1,406 narration MP3s, with 88 finished chapters live at papermorph.diamonddoge.org.
  • Verified: In-browser quizzes work exactly as described. The engine exposes quiz() with choice, tap and blanks builders and client-side grading, and site/elementary-algebra/ch01/index.html ships a working instance.
  • Verified: The README's "75-second demo" figure is accurate. The preview MP4 measures 74.94 seconds at 1600x900, H.264 with AAC audio.
  • Correction: The video presents "no image generation models" as the most surprising feature. The README lists it as a current shortfall under "Today: Opus 5.5 only. No image models, multilingual support, or BGM yet." and names image models under "Planned". It is a roadmap gap, not a design choice to celebrate.
  • Correction: The "voice narration" does not come from the AI model. scripts/tts.py synthesizes every clip with Microsoft Edge TTS via the edge-tts package. No current Claude model emits audio, per Anthropic's live models overview, so the voice has to come from outside the model.
  • Partial correction: "You just feed it a book" skips the user approvals built into the skill. SKILL.md requires intake decisions, a confirmed chapter map, a pilot chapter delivered "for the user's review", and leaves "overall aesthetics, pacing and teaching effectiveness" to a human.
  • Partial correction: The video's link reads github.com/dozentwelve/papermorph. GitHub resolves URLs case-insensitively so it works, but the canonical spelling is DozenTwelve/Papermorph, and the install command on screen, npx skills add DozenTwelve/Papermorph, needs that exact casing.
  • Unverified: Who built it. The GitHub account DozenTwelve carries the display name "TedKaczynski" and the bio "I am a beautiful person with a charming personality.", created 2017-03-15, with 6 public repositories and 8 followers. That is a pseudonym. I found no profile, blog or company field to attribute the work to a real person.
  • Context: The skill is free, the model is not. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, and the pipeline generates a 22 KB to 27 KB HTML file per chapter plus narration for each beat.
  • Context: Four of the six books on the demo bookshelf do not exist yet. The site marks physics, children's picture books and curious books "Coming soon" and states "Books on the future shelves are previews, not available lessons."

Resources

  • DozenTwelve/Papermorph on GitHub. The primary source for the repository's license, creation date, star count, file tree and all README claims quoted above.
  • Papermorph SKILL.md. Defines the six-stage pipeline, the per-chapter loop, the user approval checkpoints and the uv, ffmpeg, Edge TTS and Playwright Chromium prerequisites.
  • scripts/tts.py. Establishes that narration audio is synthesized by Microsoft Edge TTS through the edge-tts package, not by the language model.
  • references/engine.md. Documents the quiz() API, the choice, tap and blanks builders and the client-side api.grade() call behind the in-browser quizzes.
  • Papermorph live bookshelf. Confirms the 68-chapter and 20-chapter counts and labels the physics, picture-book and curious-book shelves "Coming soon".
  • Anthropic models overview. Confirms Claude Opus 5.5 is current and recommended as the default, lists its $4 and $20 per million token pricing, and states that all current models produce text output only.

Published October 5, 2026. Writeup generated from a favorited TikTok.