<- all tokdocs

Claude Code's Cache Timer Reads 60 Minutes, and Anthropic's API Default Is Five

Watch on TikTok

View on TikTok ->

Anthropic's documented default prompt cache TTL is five minutes, and the one-hour TTL is an opt-in that Claude Code requests for the main conversation only on a Claude subscription inside your plan's included usage. KodeKloud's 105-second explainer walks through why Claude Code re-sends your whole conversation on every turn, what the server keeps between turns, and what the clock in the composer counts down. The mechanism it describes matches Anthropic's own documentation. Three of the claims attached to that mechanism need correcting: who actually gets the 60-minute timer, what writing to the cache costs, and whether Anthropic describes the stored state the way the video's diagram labels it.

The mechanism is right, and Anthropic documents it in the same terms

The video's core explanation holds up line for line against the Claude Code docs. The narrator says the model remembers nothing and Claude Code sends the whole conversation again every turn. Anthropic's prompt caching page for Claude Code says the same thing: "The model doesn't remember anything between requests, so Claude Code re-sends the full context: the system prompt, your project context, every prior message and tool result, and your new message."

The prefix-match behavior the video describes is also documented. The narrator says the server "looks at the beginning of conversation and checks if it's the exact same bytes it saw a minute ago." Anthropic's wording: "The API caches by matching the start of each request, called the prefix, against content it recently processed... The match is exact, so a change anywhere in the prefix recomputes everything after it. There is no per-file or per-segment caching."

The closing claim is correct too. When the timer runs out, nothing breaks. The next request reprocesses the full history as uncached input, and Anthropic's docs describe this as "a one-time slower, more expensive turn, after which the new prefix is cached."

Two mechanics the video does not cover, both from the API reference. A request can carry at most four explicit cache_control breakpoints. Prefixes shorter than a model-specific minimum silently fail to cache with no error: 512 tokens on Opus 5.5, Opus 5, Sonnet 5.5, and Fable 5.1; 1,024 on Opus 4.8 and Sonnet 5; 2,048 on Opus 4.7 and Haiku 3.5; 4,096 on Opus 4.6 and Haiku 4.5.

The 60-minute timer depends on how you pay, and most API users get five minutes

The video says "By default, your state is stored for 5 minutes in KV cache, but Claude Code is now giving us a 1 hour cache." The first half is right. The second half is right for one group of users and wrong for everyone else.

Anthropic's Claude Code docs set the default TTL per request bucket, and the bucket's default depends on your billing path:

Request bucket Claude subscription, within plan usage Usage credits, API key, or cloud provider
Main conversation One hour Five minutes
Everything else (subagents, workflows, forks, compaction, session titles) Five minutes, except server-controlled helper requests Five minutes

So the "60m" chip the video mocks up in the composer is what a subscription user inside their plan limits sees. Sign in with an API key, use a cloud provider, or spill over into usage credits, and Claude Code drops the main conversation to the five-minute TTL. The docs are explicit about the spillover case: "Once you go over your plan's usage limit and Claude Code draws on usage credits, you are billed for that usage, so Claude Code drops the main conversation to the cheaper five-minute TTL."

You can override this. The promptCacheTtl setting or the CLAUDE_CODE_PROMPT_CACHE_TTL environment variable accepts 5m or 1h for the main conversation, and subagentPromptCacheTtl or CLAUDE_CODE_SUBAGENT_PROMPT_CACHE_TTL covers everything else. Both require Claude Code v2.1.242 or later. At the API level, the one-hour cache is a ttl field inside cache_control, and Anthropic's docs list it as generally available rather than a beta.

One more detail the video skips: a cache read refreshes the timer at no extra cost. Anthropic's pricing page labels that column "Cache hits and refreshes." Keep working and the five-minute window never expires.

Writing to the cache costs more than sending the tokens uncached

This is the claim that points the wrong way. The video says: "Claude still charges you for the tokens stored in cache, but they are 10 to 40 times cheaper than normal input tokens that you send through chat." Storing and reading are separate billing events with multipliers that move in opposite directions.

From Anthropic's pricing page:

Cache operation Multiplier on base input price
5-minute cache write 1.25x
1-hour cache write 2x
Cache read (hit or refresh) 0.1x, or 0.05x on Opus 5.5, or 0.025x on Fable 5.1 and Mythos 5.1

Writing costs 25% more than sending the same tokens uncached on the five-minute TTL, and double on the one-hour TTL. Only the read is cheap.

Here is what that means on Opus 5.5, the model the video's mockup names, at $4 per million input tokens and $0.20 per million cached tokens. Take the 100,000-token conversation the video puts on screen:

  • Sending it uncached: $0.40
  • Writing it to the five-minute cache: $0.50
  • Each later read of that same prefix: $0.02
  • Writing it to the one-hour cache: $0.80

Anthropic states the break-even directly: caching pays off "after one cache read for the 5-minute duration (1.25x write), or after two cache reads for the 1-hour duration (2x write)." The one-hour TTL earns its doubled write price only when your gaps between turns regularly exceed five minutes.

"10 to 40 times cheaper" is correct for reads, and the figure is model-specific

The video's on-screen "10-40x CHEAPER" graphic, with the orange INPUT bar next to the short purple CACHED bar, matches Anthropic's published multipliers once you confine it to cache reads. The spread comes from per-model rates rather than from any variability within one model:

Model Base input Cache read How much cheaper
Claude Fable 5.1, Mythos 5.1 $10 / MTok $0.25 / MTok 40x
Claude Opus 5.5 $4 / MTok $0.20 / MTok 20x
Claude Opus 5, Opus 4.8 $5 / MTok $0.50 / MTok 10x
Claude Sonnet 5.5, Sonnet 5 $2 / MTok $0.20 / MTok 10x
Claude Haiku 4.5 $1 / MTok $0.10 / MTok 10x

For a Claude Code session on Opus 5.5, the number is 20x, not a range. The 40x end of the video's figure applies only to Fable 5.1 and Mythos 5.1.

Anthropic does not call it a KV cache, and the state is not always on Anthropic's side

Two framing choices in the video go past what the documentation says.

First, the terminology. The video's diagram puts a purple box labeled "SAVED KV CACHE" under the Anthropic logo, and the narration describes "an internal working state called a KV cache." That is a reasonable description of how transformer inference works in general. Anthropic's prompt caching documentation does not use the term. It describes prefixes and cache entries and leaves the storage format unspecified. Treat the KV framing as the creator's explanation of the underlying mechanism rather than a documented product detail.

Second, the location. The video says "Anthropic saves this state on their side." That is true for an API key or a Claude subscription, and false for several other paths. Anthropic's Claude Code docs list where the cache lives by authentication method: Anthropic's infrastructure for an API key, a Claude subscription, or Claude Platform on AWS; your cloud provider's serving infrastructure on Amazon Bedrock and Google Cloud's Agent Platform; and Azure infrastructure for Microsoft Foundry deployments hosted on Azure. Behind a custom ANTHROPIC_BASE_URL or an LLM gateway, whether caching works at all depends on whether the gateway forwards the cache_control markers unchanged.

What the frames actually show

The 53 sampled frames are a single animated diagram, no talking head, no real terminal. The composer mockup carries three readable chips across nearly every frame: a clock icon reading 60m, a model chip reading Opus 5.5 (1M) Medium, and Bypass permissions on the right. Frame 5 shows the tooltip the video is built around: "Prompt cache warm, about 60 min left."

Opus 5.5 is a real model. It appears in Anthropic's current pricing table and in the example status line payload in the Claude Code docs as "id": "claude-opus-5-5". The video's mockup is using a current model name rather than inventing one.

The diagram's own clock contradicts its composer chip, and that tension is the point of the ending. At frame 47, the "SAVED KV CACHE" box at Anthropic carries a clock face labeled 5 MIN while the composer above it still reads 60m. Frame 53 completes the sequence: the cache box empties to dashed outlines, the GPU's token bars turn orange, and the caption reads FULL PRICE.

If you want the real version of that timer rather than a mockup, Claude Code exposes it. Since v2.1.251, the status line JSON carries a prompt_cache object with warm, ttl, expires_at, requests, misses, hit_ratio, cache_write_tokens, and last_miss_cause. Running /usage prints a Prompt cache (main) line with the session's hit ratio, miss count, and warm state, and since v2.1.260 it names the likely cause of the last miss, such as likely cause: tool definitions changed.

Key Takeaways

  • Anthropic's default prompt cache TTL is five minutes. The one-hour TTL is an opt-in ttl: "1h" field inside cache_control, generally available rather than beta.
  • Correction to the video: the 60-minute timer is not Claude Code's blanket default. Claude Code requests the one-hour TTL for the main conversation only on a Claude subscription within included plan usage. API key, cloud provider, and usage-credit sessions get five minutes, as do subagents, workflows, forks, and compaction on every plan.
  • Correction to the video: writing to the cache costs more than uncached input, at 1.25x base input price for the five-minute TTL and 2x for the one-hour TTL. Only the read is discounted.
  • The video's "10 to 40 times cheaper" holds for cache reads and is model-specific: 0.1x on most models, 0.05x on Opus 5.5 (20x cheaper), and 0.025x on Fable 5.1 and Mythos 5.1 (40x cheaper).
  • Caching breaks even after one cache read on the five-minute TTL and after two reads on the one-hour TTL, per Anthropic's pricing page.
  • A cache read refreshes the entry's timer at no extra cost, so continuous work keeps a five-minute cache warm indefinitely.
  • Correction to the video: Anthropic's documentation does not use the term "KV cache" for what it stores. It describes prefixes and cache entries without specifying the storage format.
  • Correction to the video: the cached state does not always sit on Anthropic's infrastructure. On Amazon Bedrock and Google Cloud's Agent Platform it lives in your cloud provider's serving infrastructure, and on Microsoft Foundry deployments hosted on Azure it lives on Azure.
  • Caching is prefix-based and the match is exact. A request allows at most four explicit cache_control breakpoints, and prefixes below a model-specific minimum (512 to 4,096 tokens) silently fail to cache.
  • promptCacheTtl and CLAUDE_CODE_PROMPT_CACHE_TTL let you pick 5m or 1h for the main conversation. subagentPromptCacheTtl covers everything else. Both need Claude Code v2.1.242 or later.
  • Claude Code reports live cache state in the status line prompt_cache object (v2.1.251+) and in the Prompt cache (main) line from /usage, which names the likely miss cause from v2.1.260.

Resources

Published October 2, 2026. Writeup generated from a favorited TikTok.