Is Claude Opus 5.5 actually cheaper than Claude Opus 5?
AI Architect
Key takeaways
- Claude Opus 5.5 prices input at $4 and output at $20 per million tokens, 20 percent under Claude Opus 5's $5 and $25.
- Cache reads fall from $0.50 to $0.20 per million tokens, a 60 percent cut, because the multiplier moved from 0.1x to 0.05x of base input.
- Anthropic documents that Claude Opus 5.5 tends to think more per turn than Claude Opus 5 at the same effort level, most of all at xhigh and max.
- Output tokens can rise 25 percent before a 20 percent output price cut is cancelled, so the tolerance depends on how much of the bill is output.
- A cache-heavy agent turn absorbs 2.09 times the output tokens before breaking even, while an uncached chat turn absorbs only 1.35 times.
Claude Opus 5.5 launched on 22 September 2026 at $4 and $20 per million input and output tokens, against $5 and $25 for Claude Opus 5. A flat 20 percent cut, and the first thing most teams will do is switch the model string.
Two things complicate reading the resulting bill. One cut is much larger than 20 percent, and one of the changes shipping alongside can cancel the rest of it.
What does Claude Opus 5.5 cost against Claude Opus 5?
Anthropic’s pricing page gives five rates per model. On Claude Opus 5.5 they are $4 base input, $5 for a five-minute cache write, $8 for a one-hour cache write, $0.20 for a cache read and $20 output. Claude Opus 5 is $5, $6.25, $10, $0.50 and $25.
Four of those five fall by exactly 20 percent. The cache read falls by 60 percent, and the reason is in a footnote: cache hits on Claude Opus 5.5 are priced at 0.05x the base input price, where every model outside the Fable and Mythos lines uses 0.1x.
Batch pricing halves both directions as usual, at $2 and $10. Fast mode runs $8 and $40, exactly twice standard, and Claude Opus 5’s fast mode of $10 and $50 is also twice its standard, so the premium multiplier did not move.
Why is Claude Opus 5.5’s cache read cut bigger than 20 percent?
Because two things changed at once and they compound.
A cache read price is the base input price times a multiplier. Anthropic cut the base from $5 to $4 and cut the multiplier from 0.1x to 0.05x, which gives $0.50 down to $0.20. The 60 percent fall is 0.8 multiplied by 0.5.
Which Claude Opus 5.5 price cuts behaviour cannot undo
Show as text
| # | Layer | Note |
|---|---|---|
| 1 | Cache reads, $0.50 to $0.20 per million | Multiplier moved from 0.1x to 0.05x. |
| 2 | Batch pricing still halves both rates | $2 and $10 per million on Opus 5.5. |
| 3 | Fast mode premium stays at 2x standard | $8 and $40, on the Claude API only. |
| · | fixed above, behaviour below (breakpoint) | |
| 4 | Input price, $5 to $4 per million | Undone if a request carries more state. |
| 5 | Output price, $25 to $20 per million | Undone by 25 percent more thinking. |
| 6 | Effort default drops from high to medium | Cheaper and shallower in one change. |
This matters more than the headline for anything with a long prefix. In a well-cached agent loop most input arrives as cache reads, so the term that dominates your input bill is the one that fell furthest. The break-even arithmetic for caching itself barely moves, since a five-minute write is still 1.25x base and pays back after one read, and a one-hour write is still 2x and pays back after two.
The minimum cacheable prefix on Claude Opus 5.5 is 512 tokens, which is the low end of the published range and makes the rate reachable on small prefixes.
How much extra thinking erases Claude Opus 5.5’s price cut?
Twenty-five percent, if output is the whole bill.
Thinking tokens bill as output. On Claude Opus 5.5 thinking cannot be switched
off: thinking: {"type": "disabled"} and a manual budget_tokens both return a
400 at every effort level, and the
effort parameter
is the only control. Anthropic’s
what’s new page
states that at the same effort setting the model tends to think more per turn
than Claude Opus 5, most of all at xhigh and max.
The arithmetic is one division. Output fell from $25 to $20, so $25 buys 1.25 times as many tokens at the new rate. Spend more than 25 percent extra output and a pure-output request costs more than it did.
Real requests carry input too, and the 20 percent input cut buys a little more headroom. How much depends on the mix, which is what the next section works out.
What does one Claude Opus 5.5 agent turn actually cost?
Two shapes, priced from the published rates.
A cache-heavy agent turn: 200,000 tokens of context of which 190,000 are cache reads, 10,000 fresh input, 4,000 output. On Claude Opus 5 that is $0.095 plus $0.05 plus $0.10, so $0.245. The identical token counts on Claude Opus 5.5 cost $0.038 plus $0.04 plus $0.08, so $0.158, a 35.5 percent saving rather than 20.
An uncached chat turn: 2,000 input, 1,000 output. Claude Opus 5 charges $0.035 and Claude Opus 5.5 charges $0.028, the flat 20 percent.
Cost of one turn on Claude Opus 5.5, as a share of Claude Opus 5
Show data
| Item | Value (percent of the Claude Opus 5 cost) | Note |
|---|---|---|
| Agent turn, identical token counts | 64.5 | 190k cache read, 10k input, 4k output. |
| Agent turn, 1.5x output tokens | 80.8 | Still cheaper. Cache reads dominate. |
| Chat turn, identical token counts | 80 | 2k input, 1k output, nothing cached. |
| Chat turn, 1.5x output tokens | 108.6 | A net increase over Claude Opus 5. |
Now grow the output by half. The agent turn reaches $0.198, still 19 percent under Claude Opus 5. The chat turn reaches $0.038, which is 8.6 percent above it.
Claude Opus 5.5 cost against Claude Opus 5 as thinking grows
- Cache-heavy agent turn
- Uncached chat turn
- Parity with Claude Opus 5
Show data
| Output tokens as a multiple of the Claude Opus 5 turn | Cache-heavy agent turn | Uncached chat turn | Parity with Claude Opus 5 |
|---|---|---|---|
| 1 | 0.64 | 0.8 | 1 |
| 1.5 | 0.81 | 1.09 | 1 |
| 2 | 0.97 | 1.37 | 1 |
| 2.5 | 1.13 | 1.66 | 1 |
| 3 | 1.3 | 1.94 | 1 |
Solving for parity gives 2.09 times the output tokens on the agent turn and 1.35 times on the chat turn. Caching is what buys the tolerance.
Should you move production traffic to Claude Opus 5.5 for the price?
My read is yes for cached agent workloads and with a measurement first for everything else, and that the price is the weaker half of the argument either way.
The cached case is comfortable. A 35.5 percent saving at identical token counts and a 2.09x tolerance on output growth means the migration is very hard to lose on cost, before counting the capability gains Anthropic claims for agentic coding.
The uncached case needs a number. A 1.35x tolerance is inside the range that
“thinks more per turn, most of all at xhigh and max” could plausibly produce,
and the
effort default dropping from high to medium
means an unchanged request is also a shallower one. Two variables moved, and
reading the bill without pinning effort tells you nothing about either.
What I would not do is treat 20 percent as the expected saving. It is the expected saving on a workload with no cache and no change in output length, and almost nothing looks like that.
Do this
Price a Claude Opus 5.5 migration before you move traffic
The trap is changing the price and the thinking depth in one step and then reading the bill as though only the price moved. These steps separate the two so the comparison means something.
Record your token mix on Claude Opus 5 first
Pull input tokens, cache read tokens, cache write tokens and output tokens per request from your own usage logs. The share arriving as cache reads is the single number that decides how large the Claude Opus 5.5 saving is, because that rate fell 60 percent against 20 percent on the base rates.
Set effort explicitly on both sides of the test
A request that omits effort ran at high on Claude Opus 5 and runs at medium on Claude Opus 5.5. Leaving it unset compares two different depths of reasoning and reads the difference as a price change. Pin the same level on both runs before you compare anything.
Remove any disabled-thinking setting before you switch
Both thinking.type disabled and thinking.type enabled with budget_tokens return a 400 on Claude Opus 5.5 at every effort level. Code carrying either one fails on the first request rather than costing you money, so this is a compile-time class of problem and worth clearing first.
Measure output tokens per completed task rather than per request
Anthropic documents more thinking per turn at a given effort, and a turn that thinks longer may also finish in fewer turns. Counting tokens per request can show a rise while cost per finished task falls. Define the task boundary before you start counting.
Compute your own breakeven multiple
Divide the fixed part of the Claude Opus 5 bill that Claude Opus 5.5 also charges into the total, then solve for the output tokens that reach parity. A cache-heavy turn in the worked example tolerates 2.09 times the output and an uncached one only 1.35 times.
Re-run your effort sweep rather than carrying a level over
Anthropic's prompting guidance says to re-run the sweep because the model thinks more per turn at the same setting, especially at xhigh and max. The level that was correct on Claude Opus 5 is a different amount of spend here, and one level down may hold quality at a lower bill.
Frequently asked questions
- How much cheaper is Claude Opus 5.5 than Claude Opus 5?
- Twenty percent on both input and output tokens, at $4 and $20 per million against $5 and $25. Cache reads fall further, from $0.50 to $0.20 per million, because Anthropic changed the cache read multiplier on this model from the standard 0.1x of base input to 0.05x. Batch pricing halves both rates as usual.
- Can Claude Opus 5.5 end up costing more than Claude Opus 5?
- Yes, on output-heavy requests. Thinking tokens bill as output, thinking cannot be disabled on this model, and Anthropic documents more thinking per turn at a given effort level. On an uncached request where output dominates, 1.35 times the output tokens cancels the whole cut and anything above that is a net increase.
- Does prompt caching change the Claude Opus 5.5 comparison?
- Substantially. The cache read cut is 60 percent against 20 percent on the base rates, so the more of your input arrives as cache reads the larger the saving and the more extra thinking you can absorb before breaking even. A long agent loop is the shape that benefits most.
- Why can't I just disable thinking to control Claude Opus 5.5 cost?
- Because the API rejects it. Both thinking.type disabled and a manual budget_tokens value return a 400 invalid_request_error on this model. The effort parameter is the only control, and its default on Claude Opus 5.5 is medium rather than the high that Claude Opus 5 used.
- What does fast mode cost on Claude Opus 5.5?
- $8 per million input tokens and $40 per million output, which is exactly twice the standard rate. Claude Opus 5's fast mode is $10 and $50, also twice its standard rate, so the premium multiplier is unchanged and only the absolute price moved. Fast mode is Claude API only and excluded from the Batch API.