<- all tokdocs

Scrapling's "700x Faster" Headline Only Holds Against BeautifulSoup. Against Scrapy It Is 1.035x

Watch on TikTok

View on TikTok ->

The video opens on a card reading "700x Faster Web Scraping" and then scrolls through the actual Scrapling GitHub README, including the benchmark table that the claim comes from. That table is on screen long enough to check. Scrapling's own benchmark puts its parser at 1.035x the speed of Parsel/Scrapy and 1.286x the speed of raw lxml; the 700x-class multiples in the video only appear in the rows for MechanicalSoup and BeautifulSoup, which are not what production scrapers use. Several other claims in the 68 seconds are also off: the author's name is misspelled, the star count is understated by about 11,000, and the project is two years old rather than new.

The Benchmark Table Says Three Different Things Depending On The Row

The README's "Text Extraction Speed Test (5000 nested elements)" ranks eight libraries. Scrapling is first at 1.99ms. Parsel/Scrapy is second at 2.06ms, a 1.035x gap. Raw lxml is third at 2.56ms, 1.286x. PyQuery is 23.98ms (12x), Selectolax 197.02ms (99x), MechanicalSoup 1545.15ms (776.5x), BS4 with lxml 1562.1ms (785.0x), and BS4 with html5lib 3412.73ms (~1714.9x). The README says these are averages of 100+ runs and links the methodology in benchmarks.py.

So the video's two numbers are not inconsistent with each other, which is what they look like on a first listen. They are two different rows. "700x" rounds down from the MechanicalSoup and BS4-with-lxml rows. "over 1,700 times longer" is the BS4-with-html5lib row exactly. The video just presents the first as a blanket claim about "other web scraping tools" instead of a claim about one slow parser backend. If you are already on Scrapy or calling lxml directly, Scrapling's parser buys you a few percent, not three orders of magnitude.

The ~2ms parser figure the video quotes is real and correctly stated. It is the 1.99ms Scrapling row. The caveat is that it measures one task, text extraction across 5000 nested elements, on a synthetic document. There is a second table for element similarity and text search, where Scrapling runs 2.3ms against AutoScraper's 12.58ms, a 5.47x gap. That second table is the one that backs the adaptive feature, and the video never mentions it.

The Author Is Karim Shoair, Not Kareem Shoaib

The video says "Kareem Shoaib." The repository is github.com/D4Vinci/Scrapling, the GitHub account's display name is "Karim shoair," the PyPI author field reads "Karim Shoair," and the BSD-3-Clause license header visible on screen in the video itself reads "Copyright (c) 2024, Karim shoair." The video shows the correct spelling in frame and says the wrong one out loud.

"Solo developer" is close but not exact. The commit history is dominated by D4Vinci, but the repo lists 31 contributors. It is one person's project with outside patches, not a one-man repo.

"Just Launched" Is Wrong By Two Years

The GitHub repository was created on 2024-10-13 and the first PyPI release, v0.1, went out the same day. The current release is v0.4.15, published 2026-08-23, and the project ships roughly every one to three weeks. The repo was last pushed on 2026-09-30, the day the video went up. This is a mature, actively maintained project still sitting below 1.0, not a launch.

The star count is also understated. The video says "over 74,000." The GitHub API returns 84,892 stars as of 2026-10-01, with 8,689 forks. The direction of the error is in the project's favor, which is unusual for this genre of video, and probably just means the script was written against a number that was weeks stale.

The BSD-3-Clause license claim checks out. The license tab is visible in the video at the 33-second mark and the repo's SPDX identifier is BSD-3-Clause.

The Adaptive And Stealth Features Are Real, With Narrower APIs Than Described

The video describes element relocation as "Scrapling remembers exactly what that element looked like and finds it again." The actual mechanism is two flags. You pass auto_save=True on first selection, which captures the element's tag, text, attributes, siblings, parent, and path into a SQLite store keyed by domain and an identifier. Later, when the page structure has changed, you pass adaptive=True on the same selector and Scrapling scores candidates by similarity rather than matching exactly. The docs state a limitation the video does not: only the properties of the first element in a selection result are saved, so multi-element selectors relocate one match, not the set.

For anti-bot, the video says stealth mode "clears Cloudflare's toughest challenges with a single flag." The real surface is StealthyFetcher (or StealthySession for persistent sessions) with solve_cloudflare=True. The docs list JavaScript managed challenges, interactive click-box challenges, and invisible background verification, plus custom pages with embedded captcha. The docs also recommend raising the timeout to at least 60 seconds when solving is enabled, which is the part that gets left out of a 68-second video. Fingerprint handling is separate and granular: CDP and WebRTC leak suppression, canvas noise, headless-detection patching, and block_webrtc / hide_canvas / allow_webgl parameters. The three fetcher classes are Fetcher for HTTP with TLS impersonation, DynamicFetcher for Playwright-driven browser automation, and StealthyFetcher for the anti-bot path.

The MCP Server Exists And Does More Than The Video Says

This is the one claim where the video undersells. Scrapling ships an MCP server with thirteen tools, split into one-shot tools that open and close their own browser or client per call (make_request, bulk_get, fetch, bulk_fetch, stealthy_fetch, bulk_stealthy_fetch) and session tools that hold a connection open (open_session, open_request_session, session_fetch, session_make_request, close_session, list_sessions, screenshot). Install is pip install "scrapling[ai]" then scrapling install, and you run it with scrapling-mcp, or scrapling-mcp --http for HTTP transport.

The token-cost claim in the video is the real selling point. You pass a CSS selector to narrow the page before any of it reaches the model, so the agent reads the product grid instead of the whole document. There is also a prompt-injection defense the video skips entirely: with main_content_only enabled, which is the default, the server strips CSS-hidden elements, accessibility-hidden content, template tags, HTML comments, and zero-width characters before the content reaches the model. That matters more than the token savings if you are pointing an agent at pages you do not control.

Key Takeaways

  • The "700x faster" headline is a BeautifulSoup comparison. Scrapling's own README puts it at 1.035x against Parsel/Scrapy and 1.286x against raw lxml on the same test.
  • The author is Karim Shoair (GitHub D4Vinci), not "Kareem Shoaib." The video shows the correct spelling on screen while saying the wrong one.
  • 84,892 stars as of 2026-10-01, not 74,000. The repo has 31 contributors, not one.
  • Not a launch. Created 2024-10-13, currently on v0.4.15 released 2026-08-23, still pre-1.0.
  • The ~2ms parser figure (1.99ms) and the BSD-3-Clause license are both accurate as stated.
  • The real API is auto_save=True then adaptive=True for element relocation, and StealthyFetcher(solve_cloudflare=True) for Cloudflare, with a recommended 60-second timeout.
  • The MCP server is the most underrated part: 13 tools, CSS-selector narrowing before content hits the model, and prompt-injection stripping on by default.

Resources

  • Scrapling on GitHub: the repo shown in the video; source for the benchmark tables, feature list, license, and star count.
  • Performance benchmark methodology (benchmarks.py): the script behind the 1.99ms and ~785x figures, linked from the README as the stated methodology.
  • Scrapling on PyPI: source for the author name, v0.4.15, the 2026-08-23 release date, the 2024-10-13 first release, and the Python 3.10+ requirement.
  • MCP Server guide: the thirteen tools, install commands, CSS-selector narrowing, and the main_content_only prompt-injection sanitizing.
  • StealthyFetcher docs: solve_cloudflare, the three challenge types, fingerprint-spoofing parameters, and the 60-second timeout recommendation.
  • Adaptive element relocation docs: adaptive=True, auto_save=True, adaptive_domain, the SQLite backend, and the first-element-only limitation.

Published September 30, 2026. Writeup generated from a favorited TikTok.