Prompt Caching Economics Compared Across Providers

Prompt caching is the cheapest cost lever most LLM teams have not fully exploited: providers charge a fraction of the normal input rate for any prompt prefix your workload reuses, and the discounts are large enough to reorder architecture decisions. DeepSeek bills disk-cache hits at $0.014 per million tokens against $0.14 per million for misses, cutting API costs by up to 90% when prefixes repeat. Anthropic and OpenAI both charge a 25 percent premium to write a cache entry and then a tenth of the base rate to read it back, while Google’s Gemini API applies implicit caching by default. Whether that math saves money depends on repeat rate, cache lifetime and prompt ordering — three things your code controls.

The mechanics are identical everywhere: the provider hashes your prompt prefix, stores the computed attention state, and reuses it when a later request begins with the same bytes. Matching is strictly prefix-based, so a single changed token near the front invalidates everything behind it. That constraint — stability at the front, variance at the back — separates workloads that save most of their input bill from workloads that pay write premiums for nothing.

How Providers Price Cached Tokens

Each vendor exposes the same primitive with a different pricing shape. Anthropic charges 1.25 times the base input rate to write a 5-minute cache entry, 2 times for the 1-hour variant, and 0.1 times for every read on most Claude models — Claude Opus 5.5 reads at 0.05 times base under a per-model exception. OpenAI mirrors the write premium on its newer models. Implicit caching is enabled by default on Gemini 2.5 and newer models, with a minimum cacheable prefix of 2,048 tokens on Gemini 2.5 Flash and Pro and 4,096 tokens on newer Flash generations, so short prompts never enter the cache at all. DeepSeek has no write cost and no API surface: the cache is a side effect of its disk layer.

Published rates, per million tokens:

ProviderCache writeCache read or hitActivation
Anthropic (Claude)1.25× base (5-minute); 2× (1-hour)0.1× base (0.05× on Opus 5.5)Explicit or automatic cache_control
OpenAI (GPT-5.6 and later)1.25× standard input0.1× (0.05× on GPT-6.1 Sol)Enabled by default
Google Gemini (current Flash)No write premium; hourly storage fee$0.075 cached vs $0.75 inputImplicit by default
DeepSeekNone$0.014 hit vs $0.14 missAutomatic disk cache

The Write Premium Math

The premium only pays off on reuse. On GPT-5.6 and later models, cache writes cost 1.25 times the standard input rate and reads cost 0.1 times, so writing a prefix once and fully reusing it once costs 1.35 times its ordinary input cost, versus 2 times for processing it twice uncached. One reuse already wins, narrowly; the second reuse puts you clearly ahead. Model the prefix cost per request as: effective input cost = W + (N − 1) × R, where W is the write-inclusive first pass, R the read rate and N the number of requests sharing the prefix.

Run the numbers on a concrete case. A 100,000-token system prompt, tool manifest and retrieved document set reused ten times per hour costs one write at 1.25× base plus nine reads at 0.1× base — about 2.15 times base in total, versus 10 times base uncached. That is close to an 80 percent reduction on that prefix, before touching model choice or batching. The same arithmetic collapses if the prefix mutates: every changed byte near the front converts a read into a full-price write plus a full-price read.

TTLs Decide Whether Hits Happen

Pricing is half the economics; expiry is the other half. Anthropic’s default cache lives 5 minutes and is refreshed for free each time the entry is read, with a paid 1-hour option for slower access patterns — and the clock starts at the beginning of the request, not the end of the response, so long generations eat into the window. OpenAI exposes retention controls with model-dependent values and, critically, does not share caches across organizations or regional processing boundaries, which matters for EU teams routing through European data residency. Explicit Gemini caches add $0.50 per million tokens per hour of storage, while cached input on the current Flash generation lists at $0.075 per million tokens against $0.75 uncached, so sparse reuse can cost more in storage than it saves in reads. DeepSeek needs no configuration: hits appear automatically, and the company reports first-token latency on a 128K prompt dropping from 13 seconds to 500 ms when the prefix is warm.

Break-Even: When Caching Wins

Before wiring cache markers through your stack, run this decision procedure:

  1. Measure prefix repetition. Log a hash of the opening tokens of every request and count how many requests share each prefix inside your provider’s TTL window. Below two reuses per write, caching is noise.
  2. Reorder prompts. Move system instructions, tool definitions and reference documents to the front; put timestamps, user identifiers and per-request variables at the end. Any early mutation invalidates the entire prefix.
  3. Match TTL to traffic. Interactive traffic fits the 5-minute window; batch and scheduled workloads need the 1-hour write at 2× base, or a warm-up request fired immediately before the batch.
  4. Instrument the usage fields. Read cache_read_input_tokens on Anthropic, the cached-token breakdown in the usage object on OpenAI, total_cached_tokens on Gemini and prompt_cache_hit_tokens on DeepSeek per request, and alert when the hit rate drops — a prompt refactor can silently zero your hits.

Caching also compounds with the other cost levers. It pairs naturally with model routing: cache the long shared prefix on the cheap tier and escalate only the reasoning-heavy suffix to a frontier model. And because cache behavior changes whenever you change models or tool schemas, treat hit-rate metrics like any other regression signal and wire them into your evaluation pipeline in CI so a refactor that breaks prefix stability fails before it ships.

Sources