---
title: Is Claude Opus 5.5 actually cheaper than Claude Opus 5?
url: https://deepthinkingai.org/claude-opus-5-5-price-cut/
published: 2026-09-28
author: Shekhar Singh
topic: Models & Benchmarks
tags: claude-opus-5-5, cost, prompt-caching, thinking, migration
site: DeepThinking AI
---

# Is Claude Opus 5.5 actually cheaper than Claude Opus 5?

**Summary:** Claude Opus 5.5 lists 20 percent below Claude Opus 5 on input and output, at $4 and $20 per million tokens. Its cache reads fall 60 percent, to $0.20, because the multiplier changed from 0.1x to 0.05x. Thinking cannot be disabled, and 25 percent more output tokens erases the output cut.

## Key takeaways
- Claude Opus 5.5 prices input at $4 and output at $20 per million tokens, 20 percent under Claude Opus 5's $5 and $25.
- Cache reads fall from $0.50 to $0.20 per million tokens, a 60 percent cut, because the multiplier moved from 0.1x to 0.05x of base input.
- Anthropic documents that Claude Opus 5.5 tends to think more per turn than Claude Opus 5 at the same effort level, most of all at xhigh and max.
- Output tokens can rise 25 percent before a 20 percent output price cut is cancelled, so the tolerance depends on how much of the bill is output.
- A cache-heavy agent turn absorbs 2.09 times the output tokens before breaking even, while an uncached chat turn absorbs only 1.35 times.

Claude Opus 5.5 launched on 22 September 2026 at $4 and $20 per million input and
output tokens, against $5 and $25 for Claude Opus 5. A flat 20 percent cut, and
the first thing most teams will do is switch the model string.

Two things complicate reading the resulting bill. One cut is much larger than 20
percent, and one of the changes shipping alongside can cancel the rest of it.

## What does Claude Opus 5.5 cost against Claude Opus 5?

Anthropic's [pricing page](https://platform.claude.com/docs/en/about-claude/pricing)
gives five rates per model. On Claude Opus 5.5 they are $4 base input, $5 for a
five-minute cache write, $8 for a one-hour cache write, $0.20 for a cache read
and $20 output. Claude Opus 5 is $5, $6.25, $10, $0.50 and $25.

Four of those five fall by exactly 20 percent. The cache read falls by 60
percent, and the reason is in a footnote: cache hits on Claude Opus 5.5 are
priced at 0.05x the base input price, where every model outside the Fable and
Mythos lines uses 0.1x.

Batch pricing halves both directions as usual, at $2 and $10. Fast mode runs
$8 and $40, exactly twice standard, and Claude Opus 5's fast mode of $10 and $50
is also twice its standard, so the premium multiplier did not move.

## Why is Claude Opus 5.5's cache read cut bigger than 20 percent?

Because two things changed at once and they compound.

A cache read price is the base input price times a multiplier. Anthropic cut the
base from $5 to $4 and cut the multiplier from 0.1x to 0.05x, which gives
$0.50 down to $0.20. The 60 percent fall is 0.8 multiplied by 0.5.

**Which Claude Opus 5.5 price cuts behaviour cannot undo**

1. Cache reads, $0.50 to $0.20 per million
   Multiplier moved from 0.1x to 0.05x.
2. Batch pricing still halves both rates
   $2 and $10 per million on Opus 5.5.
3. Fast mode premium stays at 2x standard
   $8 and $40, on the Claude API only.
--- fixed above, behaviour below ---
4. Input price, $5 to $4 per million
   Undone if a request carries more state.
5. Output price, $25 to $20 per million
   Undone by 25 percent more thinking.
6. Effort default drops from high to medium
   Cheaper and shallower in one change.

The three cuts above the line are properties of the price sheet and hold whatever the model does. The three below it are cuts per token, and the number of tokens is set by behaviour that Anthropic's own documentation says has changed.

This matters more than the headline for anything with a long prefix. In a
[well-cached agent loop](/prompt-caching-economics/) most input arrives as cache
reads, so the term that dominates your input bill is the one that fell furthest.
The break-even arithmetic for caching itself barely moves, since a five-minute
write is still 1.25x base and pays back after one read, and a one-hour write is
still 2x and pays back after two.

The minimum cacheable prefix on Claude Opus 5.5 is 512 tokens, which is the low
end of the published range and makes the rate reachable on small prefixes.

## How much extra thinking erases Claude Opus 5.5's price cut?

Twenty-five percent, if output is the whole bill.

Thinking tokens bill as output. On Claude Opus 5.5 thinking cannot be switched
off: `thinking: {"type": "disabled"}` and a manual `budget_tokens` both return a
400 at every effort level, and the
[effort parameter](https://platform.claude.com/docs/en/build-with-claude/effort)
is the only control. Anthropic's
[what's new page](https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5)
states that at the same effort setting the model tends to think more per turn
than Claude Opus 5, most of all at `xhigh` and `max`.

The arithmetic is one division. Output fell from $25 to $20, so $25 buys 1.25
times as many tokens at the new rate. Spend more than 25 percent extra output and
a pure-output request costs more than it did.

Real requests carry input too, and the 20 percent input cut buys a little more
headroom. How much depends on the mix, which is what the next section works out.

## What does one Claude Opus 5.5 agent turn actually cost?

Two shapes, priced from the published rates.

A cache-heavy agent turn: 200,000 tokens of context of which 190,000 are cache
reads, 10,000 fresh input, 4,000 output. On Claude Opus 5 that is $0.095 plus
$0.05 plus $0.10, so $0.245. The identical token counts on Claude Opus 5.5 cost
$0.038 plus $0.04 plus $0.08, so $0.158, a 35.5 percent saving rather than 20.

An uncached chat turn: 2,000 input, 1,000 output. Claude Opus 5 charges $0.035
and Claude Opus 5.5 charges $0.028, the flat 20 percent.

**Cost of one turn on Claude Opus 5.5, as a share of Claude Opus 5**

| Item | Value (percent of the Claude Opus 5 cost) | Note |
|---|---|---|
| Agent turn, identical token counts | 64.5 | 190k cache read, 10k input, 4k output. |
| Agent turn, 1.5x output tokens | 80.8 | Still cheaper. Cache reads dominate. |
| Chat turn, identical token counts | 80 | 2k input, 1k output, nothing cached. |
| Chat turn, 1.5x output tokens | 108.6 | A net increase over Claude Opus 5. |

Four turns priced from the published rates. The only bar above 100 is the uncached request whose output grew by half, and it is the shape most chat products have. The agent turn tolerates the same growth because 190k of its 200k input arrives at the cache read rate.

Source: Computed from Anthropic's published per-token prices

Now grow the output by half. The agent turn reaches $0.198, still 19 percent
under Claude Opus 5. The chat turn reaches $0.038, which is 8.6 percent above it.

**Claude Opus 5.5 cost against Claude Opus 5 as thinking grows**

| Output tokens as a multiple of the Claude Opus 5 turn | Cache-heavy agent turn | Uncached chat turn | Parity with Claude Opus 5 |
|---|---|---|---|
| 1 | 0.64 | 0.8 | 1 |
| 1.5 | 0.81 | 1.09 | 1 |
| 2 | 0.97 | 1.37 | 1 |
| 2.5 | 1.13 | 1.66 | 1 |
| 3 | 1.3 | 1.94 | 1 |

Each line crosses parity at a different point. The uncached turn breaks even at 1.35 times its original output, the cache-heavy turn at 2.09 times. Both inputs are published prices, so you can substitute your own token mix and re-run it.

Source: Computed from Anthropic's published per-token prices

Solving for parity gives 2.09 times the output tokens on the agent turn and 1.35
times on the chat turn. Caching is what buys the tolerance.

## Should you move production traffic to Claude Opus 5.5 for the price?

My read is yes for cached agent workloads and with a measurement first for
everything else, and that the price is the weaker half of the argument either way.

The cached case is comfortable. A 35.5 percent saving at identical token counts
and a 2.09x tolerance on output growth means the migration is very hard to lose
on cost, before counting the capability gains Anthropic claims for agentic coding.

The uncached case needs a number. A 1.35x tolerance is inside the range that
"thinks more per turn, most of all at `xhigh` and `max`" could plausibly produce,
and the
[effort default dropping from high to medium](/claude-opus-5-5-silent-breaking-changes/)
means an unchanged request is also a shallower one. Two variables moved, and
reading the bill without pinning effort tells you nothing about either.

What I would not do is treat 20 percent as the expected saving. It is the
expected saving on a workload with no cache and no change in output length, and
almost nothing looks like that.

<ReadNext
  href="/prompt-caching-economics/"
  kicker="Related"
  title="When does prompt caching actually save money?"
  note="Where the 0.05x cache read rate sits in the break-even arithmetic, and why the multiplier barely moves it."
/>

## Price a Claude Opus 5.5 migration before you move traffic

The trap is changing the price and the thinking depth in one step and then reading the bill as though only the price moved. These steps separate the two so the comparison means something.

1. **Record your token mix on Claude Opus 5 first**: Pull input tokens, cache read tokens, cache write tokens and output tokens per request from your own usage logs. The share arriving as cache reads is the single number that decides how large the Claude Opus 5.5 saving is, because that rate fell 60 percent against 20 percent on the base rates.
2. **Set effort explicitly on both sides of the test**: A request that omits effort ran at high on Claude Opus 5 and runs at medium on Claude Opus 5.5. Leaving it unset compares two different depths of reasoning and reads the difference as a price change. Pin the same level on both runs before you compare anything.
3. **Remove any disabled-thinking setting before you switch**: Both thinking.type disabled and thinking.type enabled with budget_tokens return a 400 on Claude Opus 5.5 at every effort level. Code carrying either one fails on the first request rather than costing you money, so this is a compile-time class of problem and worth clearing first.
4. **Measure output tokens per completed task rather than per request**: Anthropic documents more thinking per turn at a given effort, and a turn that thinks longer may also finish in fewer turns. Counting tokens per request can show a rise while cost per finished task falls. Define the task boundary before you start counting.
5. **Compute your own breakeven multiple**: Divide the fixed part of the Claude Opus 5 bill that Claude Opus 5.5 also charges into the total, then solve for the output tokens that reach parity. A cache-heavy turn in the worked example tolerates 2.09 times the output and an uncached one only 1.35 times.
6. **Re-run your effort sweep rather than carrying a level over**: Anthropic's prompting guidance says to re-run the sweep because the model thinks more per turn at the same setting, especially at xhigh and max. The level that was correct on Claude Opus 5 is a different amount of spend here, and one level down may hold quality at a lower bill.


## Frequently asked questions

### How much cheaper is Claude Opus 5.5 than Claude Opus 5?

Twenty percent on both input and output tokens, at $4 and $20 per million against $5 and $25. Cache reads fall further, from $0.50 to $0.20 per million, because Anthropic changed the cache read multiplier on this model from the standard 0.1x of base input to 0.05x. Batch pricing halves both rates as usual.

### Can Claude Opus 5.5 end up costing more than Claude Opus 5?

Yes, on output-heavy requests. Thinking tokens bill as output, thinking cannot be disabled on this model, and Anthropic documents more thinking per turn at a given effort level. On an uncached request where output dominates, 1.35 times the output tokens cancels the whole cut and anything above that is a net increase.

### Does prompt caching change the Claude Opus 5.5 comparison?

Substantially. The cache read cut is 60 percent against 20 percent on the base rates, so the more of your input arrives as cache reads the larger the saving and the more extra thinking you can absorb before breaking even. A long agent loop is the shape that benefits most.

### Why can't I just disable thinking to control Claude Opus 5.5 cost?

Because the API rejects it. Both thinking.type disabled and a manual budget_tokens value return a 400 invalid_request_error on this model. The effort parameter is the only control, and its default on Claude Opus 5.5 is medium rather than the high that Claude Opus 5 used.

### What does fast mode cost on Claude Opus 5.5?

$8 per million input tokens and $40 per million output, which is exactly twice the standard rate. Claude Opus 5's fast mode is $10 and $50, also twice its standard rate, so the premium multiplier is unchanged and only the absolute price moved. Fast mode is Claude API only and excluded from the Batch API.


## Sources
- [What's new in Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5). Anthropic, 2026-09-22
- [Pricing](https://platform.claude.com/docs/en/about-claude/pricing). Anthropic
- [Effort](https://platform.claude.com/docs/en/build-with-claude/effort). Anthropic
- [Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5). Anthropic

---
Canonical HTML: https://deepthinkingai.org/claude-opus-5-5-price-cut/