Sonnet 5.5 cache reads fell from $0.20 to $0.10 per MTok on Oct 7. Break-even stays at 2 requests (5-min cache) or 3 (1-hour); each hit costs half as much.
On October 7, 2026, Anthropic cut the price of prompt-cache reads on Claude Sonnet 5.5 from $0.20 to $0.10 per million tokens, which is 0.05x the base input price instead of 0.1x. Cache writes and every other Sonnet 5.5 price are unchanged. The break-even point for caching does not move (two requests for the 5-minute cache, three for the 1-hour cache); each cache hit now costs half as much, which matters most for long agent loops.
What happened
Anthropic's release notes for October 7 say the company lowered "the price of prompt cache reads on Claude Sonnet 5.5 from $0.20 USD to $0.10 USD per million tokens: 0.05x the base input price instead of 0.1x. Cache writes and all other prices are unchanged." The pricing page now lists these Sonnet 5.5 rates, per million tokens (MTok):
| Item | Sonnet 5.5 |
|---|---|
| Base input | $2.00 |
| 5-minute cache write (1.25x) | $2.50 |
| 1-hour cache write (2x) | $4.00 |
| Cache hit or refresh (0.05x) | $0.10 |
| Output | $10.00 |
| Batch input / output (50% off) | $1.00 / $5.00 |
The pricing page also lists Claude Opus 5.5 at 0.05x ($0.20 on a $4 base) and Fable 5.1 and Mythos 5.1 at 0.025x ($0.25 on a $10 base). Sonnet 5, at the same $2/$10 base (see our Opus 5 and Sonnet price freeze coverage), still shows $0.20 for reads. Opus 5.5 was already at 0.05x before October 7 (its What's-new page lists $0.20), so Sonnet 5.5 now matches it.
Why it matters
Caching has two costs and one gain: a write premium on the first request (1.25x or 2x base input), and a discounted read on every request after it. Halving the read price does not change when you start winning, because the write premium is untouched. It raises how much you win on each reuse.
The pricing page says a Sonnet 5.5 cache hit costs 5% of the standard input price. For a prefix that is read many times, the cached portion of your input bill approaches one twentieth of the uncached price, down from one tenth.
The effect is largest where one prefix is read many times: agent loops that resend a growing conversation on every tool call. It is smallest where most of each prompt changes from request to request, which describes most retrieval-augmented generation (RAG) workloads.
Technical details
All dollar figures below are ThreatFrontier's own arithmetic from the published rates, per MTok of cached prefix, input side only. Output is billed the same with or without caching. The scenarios are illustrative assumptions, not measurements.
Break-even: unchanged at two requests (5-minute cache)
Take a prefix reused N times in a row, all inside the cache lifetime. Uncached it costs $2 per request. Cached, you pay one write plus N−1 reads.
| Requests sharing the prefix | Uncached | 5-min write, old read ($0.20) | 5-min write, new read ($0.10) | 1-hour write, new read ($0.10) |
|---|---|---|---|---|
| 1 | $2.00 | $2.50 | $2.50 | $4.00 |
| 2 | $4.00 | $2.70 | $2.60 | $4.10 |
| 3 | $6.00 | $2.90 | $2.70 | $4.20 |
| 5 | $10.00 | $3.30 | $2.90 | $4.40 |
| 10 | $20.00 | $4.30 | $3.40 | $4.90 |
| 20 | $40.00 | $6.30 | $4.40 | $5.90 |
- 5-minute cache: a second request already wins ($2.60 against $4.00), both before and after the cut. A prefix used once loses 25% to the write premium ($2.50 against $2.00).
- 1-hour cache: it loses at two requests ($4.10 against $4.00) and wins from the third, again unchanged by the cut.
- What the cut changes: the marginal cost of each extra reuse falls from $0.20 to $0.10 per MTok. At 20 requests the 5-minute total is $4.40 instead of $6.30, a 30% drop in the prefix's cost.
Two conditions come from Anthropic's prompt caching documentation. The minimum cacheable prompt for Sonnet 5.5 is 512 tokens; shorter prompts are processed without caching and no error is returned. And the cache lifetime is measured from the start of the request that wrote or read the entry, and each hit refreshes it at no extra cost. A response that streams for four minutes leaves about a minute for the next call to reuse the 5-minute entry.
Worked example 1: a 40-turn agent loop
Assume an agent starts with a 10,000-token system prompt and tool list, and every turn adds 3,000 tokens (tool results plus the model's reply) that are resent on the next call. Over 40 turns, the agent sends 2.74 MTok of input in total, if nothing is cached.
| Input-side cost for the run | Cost |
|---|---|
| No caching (2.74 MTok at $2) | $5.48 |
| Caching, reads at $0.20 (0.127 MTok written at $2.50, 2.613 MTok read) | $0.84 |
| Caching, reads at $0.10 | $0.58 |
The cut takes the cached run from $0.84 to $0.58, 31% less, and caching overall removes about 89% of the uncached input bill. The older, higher read price was already a large discount; the new one makes the remaining read cost a smaller share than the write cost ($0.26 of reads against $0.32 of writes).
40-turn agent loop, input side only. ThreatFrontier arithmetic from Anthropic's published prices.
Worked example 2: a RAG endpoint
Assume 1,000 queries inside one cache window, each with a 6,000-token static prefix (instructions and examples), 4,000 tokens of retrieved passages and a 300-token question. Only the static prefix can be cached, because the retrieved passages differ per query.
| Input-side cost for 1,000 queries | Cost |
|---|---|
| No caching (10.3 MTok at $2) | $20.60 |
| Prefix cached, reads at $0.20 | $9.81 |
| Prefix cached, reads at $0.10 | $9.21 |
Caching still cuts the bill by about 55%, but the read cut adds only 6% on top, because $8.60 of the $9.21 is the uncached passages and questions. A RAG system gains most from this change if it can keep more of the prompt stable: a fixed corpus summary, a large tool schema or a conversation history that grows in place.
Where caching still loses: sparse traffic
Take a prefix requested once every ten minutes, six times an hour. The default 5-minute entry expires before each request, so every call pays the write premium and never reads: $15.00 per MTok of prefix for six requests, against $12.00 uncached. The 1-hour entry costs $4.00 plus five reads at $0.10, or $4.50, against $12.00 uncached ($5.00 at the old read price).
Other multipliers
- Batch: Anthropic says cache multipliers stack with the Batch API discount. By our arithmetic, a Batch cache read on Sonnet 5.5 would be $0.05 per MTok; Anthropic does not publish that figure.
- US-only inference:
inference_geo: "us"applies a 1.1x multiplier to all token categories, including cache reads. - Cloud platforms: Anthropic's caching documentation says Bedrock has its own per-model minimums and failure behavior, and we did not confirm whether the new read price applies there or on Google Cloud, where partner pricing pages govern.
What teams should do
- Cache the stable prefix first. Put the system prompt, tools and fixed examples before anything volatile, and check
cache_read_input_tokensin the response usage. If it stays at zero across repeated calls, something in the prefix is changing. Our prompt caching explainer covers the ordering rule. - Count your reuse before choosing a lifetime. Three or more requests inside an hour justify the 1-hour write; a prefix hit at least every five minutes is cheaper on the default 5-minute entry.
- Re-estimate agent budgets. Long-running agents with many tool calls carry the largest cache-read volume, so they see the largest drop. Our Sonnet 5.5 efficiency guide quotes the earlier $0.20 read price; use $0.10 for current budgets. For teams still on the older model, see the Sonnet 4.5 retirement migration guide.
- Check the smaller-model option. Haiku 5.5 lists a $0.01 read price on a $0.10 base for prompts up to 100,000 tokens, and a $0.05 read price above that. Whether it is good enough for the task is a quality question these price tables cannot answer.
- Measure with your own traffic. The examples above use assumed token counts; your hit rate, prefix size and gap between calls decide the actual saving.
What is still unclear
- Whether the lower read price applies on Amazon Bedrock and Google Cloud, where those providers publish their own rates.
- Whether Anthropic will publish a combined Batch-plus-cache figure.
- Whether real workloads reach the hit rates assumed in the examples. The scenarios are ours, not Anthropic's.
Sources
- Anthropic, Claude Developer Platform release notes, October 7, 2026
- Anthropic, Pricing, prompt caching and data residency sections
- Anthropic, Prompt caching, cache lifetime and minimum cacheable length