The Second Feature Is Where Vibe-Coded SaaS Breaks, and GitClear Measured Refactoring Falling From 24.1% to 9.5% of Changed Lines
Watch on TikTok
A 50-second vertical clip at 1080x1920, posted 2026-10-02 at 15:30:53 UTC to @realcolegordon (channel name "Cole Gordon"), carrying 5,070 views, 147 likes, 30 comments, 52 saves and 4 reposts at capture, which works out to a 2.90% like rate and a 1.03% save rate. Audio is listed as "original sound" by Cole Gordon, so this is interview audio rather than a licensed track. I analyzed 25 frames sampled at 2-second intervals covering 0 to 48 seconds, plus a 231-word transcript, a pace of 277 words per minute. There is no code, no terminal, no dashboard, no screen recording and no count-up animation anywhere in the 25 frames. Every frame is the same locked-off medium shot: a blond man in a grey flannel button-down and dark grey jeans, seated in a walnut-and-black-leather lounge chair against a plain pale-grey cyclorama, speaking into a black Shure-style broadcast microphone on a boom arm. Frames 001 and 002 carry a white rounded title card reading "What Is Wrong With Vibe Coding A SaaS?" over the lower third. Frames 004 and 005 add a hand-drawn white arrow pointing up at the speaker's shoulder with the caption "CMO of $30M Company". Everything else is word-by-word burned-in captions in white sans-serif, one to four words per frame, tracking the voiceover exactly: "a software.", "The first thing", "inevitably", "the AI coding agent", "decisions.", "Somebody has to.", "does that,", "It does not know,", "intent is", "also obviously", "happen", "no context", "other than", "100% of the time", "maybe the initial", "If it does work,", "expand this thing,", "this feature set,", "Well, the original", "if you think of like", "to do those things.", "to move forward,", "Now this is common", "AI or not AI.", "a single". The camera pushes in slowly across the clip, so frames 010 onward crop progressively tighter until the chair leaves frame entirely.
The "100% of the Time" Claim Is the One Checkable Number, and It Is Overstated
Frame 014 burns "100% of the time" on screen, and the voiceover matches it: "So what will happen 100% of the time is maybe the initial thing works, maybe it doesn't." The sentence is self-undermining on its own terms. A claim that something happens 100% of the time, followed immediately by "maybe it works, maybe it doesn't," asserts certainty about an outcome and then declines to name the outcome. Nothing is being quantified.
The measured record is narrower and more interesting. The 2025 DORA report, published 23 September 2025 from a survey of nearly 5,000 technology professionals plus over 100 hours of qualitative data, found 90% of respondents using AI at work and a positive relationship between AI adoption and both software delivery throughput and product performance. The negative relationship DORA found was specifically with delivery stability: more change failures, more rework, longer time to resolve. That is a real cost, and it is not a universal failure rate.
METR's randomized controlled trial, published 10 July 2025, found 16 experienced open-source developers took 19% longer to complete 246 real issues when allowed to use AI tools, while estimating afterward that AI had made them 20% faster. METR states explicitly that this result does not show AI fails to speed up most developers generally, and that the sample was experienced contributors working in mature repositories they already knew. The clip's subject is the opposite case: a non-engineer building greenfield. METR's caveats apply, and the number does not transfer.
By Fowler's Definition, the Thing He Describes Is Not a Refactor
The voiceover says: "Now, literally, to move forward, you have to do what's called a refactor." The setup is that the original architecture cannot support new users or new features, so the foundation has to change.
Martin Fowler's definition, the one the industry actually uses, is "a change made to the internal structure of software to make it easier to understand and cheaper to modify without changing its observable behavior." Preserving observable behavior is the constraint that makes the word mean anything. Adding a feature set, adding user tiers, or rebuilding a data model to support multi-tenancy changes observable behavior by definition. What the clip describes is a re-architecture or a rewrite. The video gets the term wrong.
The distinction matters commercially. A refactor is a bounded, verifiable operation with a test suite as its oracle. A rewrite has no oracle, no fixed scope, and is the specific thing that kills shipping schedules. Calling the second one by the first one's name makes it sound cheaper than it is.
The Architecture Argument Is the Strongest Thing He Says, and DORA Backs It
Frames 004 through 007 carry the core claim: the AI coding agent makes architecture decisions, "somebody has to, it ain't going to be you." That part holds up.
DORA's 2025 finding is close to a direct restatement: teams in loosely coupled architectures with fast feedback loops see gains from AI, while teams in tightly coupled systems with slow processes see little or no benefit. AI raises change volume. Whether that volume turns into throughput or into instability is decided by the architecture it lands in. A vibe-coded first version is the least likely codebase to have loose coupling or a fast feedback loop, because nobody chose either one.
GitClear's AI code quality research, covering 211 million changed lines of code from January 2020 through December 2024, measures the mechanism. Refactored ("moved") lines fell from 24.1% of changed lines in 2020 to 9.5% in 2024. Copy-pasted lines rose from 8.3% to 12.3% over the same window, a 48% relative increase. 2024 was the first year on record where newly introduced duplicate code exceeded refactoring activity. The foundation problem the clip gestures at has a number attached to it, and the clip does not use it.
Nobody Understanding the Codebase Is a Real Failure Mode, and Developers Report It
The transcript cuts off mid-argument on the sharpest point: "The problem is now you don't have a single engineer that understands your code base." Frame 025, at 48 seconds, shows "a single" as the last caption. The clip ends there.
The 2025 Stack Overflow Developer Survey AI section measures the downstream version of this. 84% of developers are using or planning to use AI tools, up from 76% the prior year. Only 3.1% highly trust the accuracy of AI output, while 45.7% actively distrust it. The top reported frustration, at 66%, is "AI solutions that are almost right, but not quite." 45.2% say debugging AI-generated code is more time-consuming than writing it. 76% refuse to use AI for deployment and monitoring, and 4.4% think AI handles complex tasks well.
Those are professional engineers reporting a comprehension gap in code they can read. A founder who cannot read the code has the same gap with no instrument for detecting it.
Agents Making Architecture Decisions Is a Workflow Choice, Not a Property of the Tool
The clip's framing is fatalistic: the AI will make the architecture decisions, you will not, and it does not know your intent. The first half of that is accurate about what happens by default. The second half is contradicted by the vendor documentation.
Anthropic's Claude Code best practices prescribes a four-phase workflow of explore, plan, implement, commit, with plan mode as a distinct permission state where the agent reads files and answers questions without writing anything. The documented recommendation is to open the plan in an editor and edit it before approving. The same page lists "architectural decisions specific to your project" as a thing to put in a CLAUDE.md file that loads at the start of every session, and recommends having the agent interview you and write a spec to SPEC.md before any implementation begins. Its named failure pattern, "the trust-then-verify gap," is described as a plausible-looking implementation that does not handle edge cases, with the prescribed fix being "if you can't verify it, don't ship it."
So the intent gap is real, and the vendor ships a documented mechanism for closing it. The decision the clip describes as inevitable is the decision to skip that mechanism.
The On-Screen Credential Is Unsourced
Frames 004 and 005 label the speaker "CMO of $30M Company" with an arrow, and never name him. The account belongs to Cole Gordon, founder and CEO of Closers.io, whose own business is widely reported at roughly $30M in annual revenue. Gordon's title is CEO, not CMO, so the arrow is pointing at a guest. I could not identify the guest, the company, or the source of the $30M figure from any primary record. Treat the number as an unsourced on-screen caption.
Key Takeaways
- The clip's only quantified claim, "100% of the time," is immediately qualified by "maybe it works, maybe it doesn't." No outcome is actually asserted.
- Correction: the video calls the fix a "refactor." Fowler's definition requires observable behavior to stay unchanged. Adding users and feature sets changes it. The operation described is a rewrite.
- GitClear measured 211M changed lines from 2020 to 2024: refactored lines fell 24.1% to 9.5%, copy-pasted lines rose 8.3% to 12.3%, and 2024 was the first year duplicate-code introduction exceeded refactoring.
- DORA 2025 (23 Sep 2025, ~5,000 respondents) found AI adoption positively related to throughput and product performance, and negatively related to delivery stability. The clip's blanket failure framing overstates it.
- METR (10 Jul 2025): 16 experienced developers, 246 issues, 19% slower with AI while self-reporting 20% faster. METR says the result does not generalize to greenfield or non-expert work, which is the clip's actual scenario.
- Stack Overflow 2025: 84% use or plan to use AI tools, 3.1% highly trust its accuracy, 45.7% distrust it, 66% cite "almost right, but not quite" as their top frustration, 45.2% lose time debugging AI output.
- Anthropic's own docs prescribe plan mode before implementation, architectural decisions recorded in CLAUDE.md, and a written SPEC.md produced by interviewing the human. The intent gap the clip calls inevitable has a documented workflow against it.
- "Vibe coding" was coined by Andrej Karpathy in February 2025 and named Collins English Dictionary's Word of the Year for 2025 on 6 November 2025. The term is roughly 20 months old at this video's post date.
- Unverified: the identity of the speaker, the "$30M Company" he is CMO of, and the source of that revenue figure. The caption names no person and no company. Cole Gordon, who owns the account, is CEO of Closers.io rather than a CMO, so the arrow points at an unnamed guest.
Resources
- What Is Wrong with Vibe Coding a SaaS? (TikTok) — the source clip; confirms 50s runtime, 1080x1920, 2026-10-02 post date, and the "original sound" audio credit.
- Martin Fowler, Definition of Refactoring — confirms refactoring requires unchanged observable behavior, which is what makes the video's use of the term wrong.
- GitClear, AI Copilot Code Quality 2025 Research — confirms 211M changed lines analyzed, Jan 2020 to Dec 2024, and the duplication-versus-refactoring crossover.
- Report summary: GitClear AI code quality research 2025 — pins the exact year-by-year figures: moved lines 24.1% (2020) to 9.5% (2024), copy-pasted 8.3% to 12.3%.
- Announcing the 2025 DORA Report (Google Cloud) — confirms 23 Sep 2025 publication, ~5,000 respondents, 90% AI usage, the throughput-positive / stability-negative split, and the loosely-coupled-architecture finding.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — confirms 19% slowdown, 16 developers, 246 issues, 10 Jul 2025, and METR's explicit limits on generalizing the result.
- Stack Overflow 2025 Developer Survey, AI section — confirms 84% adoption, 3.1% high trust, 45.7% distrust, 66% "almost right but not quite," 45.2% debugging time loss.
- Claude Code best practices (Anthropic) — confirms the explore/plan/implement/commit workflow, plan mode, CLAUDE.md for architectural decisions, the SPEC.md interview prompt, and the "trust-then-verify gap" failure pattern.
- Vibe coding (Wikipedia) — confirms Karpathy coined the term in February 2025 and that Collins named it Word of the Year on 6 November 2025.
Published October 2, 2026. Writeup generated from a favorited TikTok.