DeepThinking AI

How do you debug Claude prompt cache misses?

AI Architect

Key takeaways

  • Claude cache diagnostics stores a fingerprint only for requests that include the diagnostics object, then compares the next request by previous response id.
  • Claude reports only the earliest divergence, so a model or system change can hide later tool or message drift until you fix the first mismatch.
  • Claude's previous_message_not_found result does not mean the prompt changed; it usually means no stored fingerprint was available to compare.
  • Claude's unavailable result covers other prompt-affecting parameters such as tool_choice, thinking, context_management, output_config, output_format, beta headers, or a conversation beyond the comparison horizon.
  • Anthropic made the response diagnostics field always present on 23 September 2026, which means parsers should distinguish null, pending comparison, and a concrete miss reason.

Claude’s cache diagnostics page answers a real pain point that the prompt caching page left open: a cache miss used to look like a single number falling to zero. On 23 September 2026 Anthropic also changed the response shape so POST /v1/messages always includes a diagnostics field, with null when you did not ask for a comparison. That sounds cosmetic. It is actually the detail that makes the feature safe to automate, because a parser can now distinguish “no diagnostics requested” from “comparison still pending” from “a divergence was found”.

What does Claude cache diagnostics actually compare?

Anthropic’s docs describe cache diagnostics as a comparison between consecutive requests, keyed by response id rather than by conversation id. When a request includes the diagnostics object, Claude stores a lightweight fingerprint for that turn. On the next request, you pass the earlier response id as diagnostics.previous_message_id, and the API compares the new request against the stored fingerprint.

The docs say the fingerprint carries hashes and token count estimates. It does not store raw prompt text. It is scoped to your organization and workspace.

The important correction is simpler. This comparison explains request structure. Claude checks whether the model, system prompt, tools, and earlier message history still line up. Cache confirmation still lives in usage.cache_read_input_tokens. If you are already using prompt caching economics to decide whether caching is worth the effort, cache diagnostics is the mechanism that tells you where that effort is leaking.

What Claude compares before naming a cache miss

What Claude compares before naming a cache missDiagram: 6 ordered layers. Model selection, then System prompt bytes, then Tools array and schemas, then First divergence here (breakpoint), then Earlier message history, then Later mismatches stay hidden.1Model selectionCache fingerprints are per model.2System prompt bytesDynamic values here force a miss.3Tools array and schemasOrder and JSON stability matter.First divergence here4Earlier message historyHistory must stay append only.5Later mismatches stay hiddenFix the first reason, then rerun.
Show as text
What Claude compares before naming a cache miss. Diagram: 6 ordered layers. Model selection, then System prompt bytes, then Tools array and schemas, then First divergence here (breakpoint), then Earlier message history, then Later mismatches stay hidden.
#LayerNote
1Model selectionCache fingerprints are per model.
2System prompt bytesDynamic values here force a miss.
3Tools array and schemasOrder and JSON stability matter.
·First divergence here (breakpoint)
4Earlier message historyHistory must stay append only.
5Later mismatches stay hiddenFix the first reason, then rerun.
Anthropic returns the earliest divergence only. That makes the result actionable, but it also means a model or system mismatch can hide tool or message drift until the earlier issue is fixed.

What does each Claude cache miss reason mean?

Anthropic reports only the earliest divergence, which is exactly the behavior you want during triage. model_changed means the model switched between turns, often because a router, A/B test, or fallback path picked another model. system_changed means the system parameter was byte different, which the docs call out as a common side effect of interpolating timestamps or request ids. tools_changed covers added, removed, or reordered tools and non-deterministically serialized input_schema JSON. messages_changed means the model, system, and tools still match, but earlier history was edited, reordered, or truncated instead of appended to.

Two results matter because they are easy to misread. previous_message_not_found means no stored fingerprint existed for the supplied response id. unavailable means Claude could not produce a comparison even though the request still ran. Anthropic names concrete causes here: other prompt-affecting parameters changed, including tool_choice, thinking, context_management, output_config, output_format, or the set of active beta headers, or the conversation drifted beyond the comparison horizon. The fix is to stabilise the request envelope, then rerun. A later problem may still be waiting behind the first one.

How a diagnosed Claude turn becomes the next comparison

How a diagnosed Claude turn becomes the next comparisonSequence diagram between Client and Claude API. 1. Client to Claude API: Send request with diagnostics enabled. 2. Claude API to Client: Return response id. 3. Client to Claude API: Send next request with previous_message_id. 4. Claude API to Client: Return diagnostics plus usage counters.ClientClaude APISend request with diagnostics enabledAPI stores a fingerprint for this turn.Return response idSave it beside the conversation.Send next request with previous_message_idSame model and stable prefix expected.Return diagnostics plus usage countersCompare structure, then confirm the hit.
Show as text
How a diagnosed Claude turn becomes the next comparison. Sequence diagram between Client and Claude API. 1. Client to Claude API: Send request with diagnostics enabled. 2. Claude API to Client: Return response id. 3. Client to Claude API: Send next request with previous_message_id. 4. Claude API to Client: Return diagnostics plus usage counters.
#FromToMessage
1ClientClaude APISend request with diagnostics enabled. API stores a fingerprint for this turn.
2Claude APIClientReturn response id. Save it beside the conversation.
3ClientClaude APISend next request with previous_message_id. Same model and stable prefix expected.
4Claude APIClientReturn diagnostics plus usage counters. Compare structure, then confirm the hit.
Cache diagnostics is a two-turn mechanism. The first diagnosed turn creates the fingerprint. The next turn supplies that response id so Anthropic can compare the new request against the stored structure.

Why can Claude say previous_message_not_found when nothing changed?

Because that result reflects missing comparison state. It says nothing yet about prompt drift. Anthropic says plainly that no fingerprint exists when the earlier request did not include the diagnostics object, came from another workspace, or is too old to retain. In other words, previous_message_not_found is often a logging or retention issue that happens before any structural comparison could begin.

That is why I would not treat this result as evidence of a cache regression in itself. The safer reading is that the debugging chain is broken. Re-enable diagnostics on every turn, keep consecutive turns close together, and make sure your application persists the previous response id beside the conversation. The same discipline matters for other cache-preserving features on Claude. A mid-conversation tool change only helps if the application keeps the surrounding request stable enough for the comparison to mean anything. The id plumbing is dull, but it is the part that turns the API from “I think we changed something” into a concrete explanation.

Why should you read Claude diagnostics with usage counters?

Because a structural comparison and a real cache hit answer different questions. Anthropic’s docs say the diagnostics field can be null when you did not ask for diagnostics, when previous_message_id was null on the first compared turn, or when a comparison found no divergence. They also document a second shape, {"cache_miss_reason": null}, for comparisons that were still running when the response was serialized. Neither shape tells you by itself whether cached tokens were actually read.

The usage counters do. usage.cache_read_input_tokens shows how many prompt tokens came from cache, and usage.cache_creation_input_tokens shows how many were written into a new cache entry. Anthropic’s prompt caching page adds one more operational trap: prompts below the model’s minimum cacheable length are processed without caching and without an error. That means you can have a stable request and still see zero read tokens. Use diagnostics to find drift. Use usage to confirm that the cached prefix existed, was long enough, and was actually reused. That division is the key insight the release note summary does not make explicit.

Should you enable Claude cache diagnostics on every cached conversation?

My view is yes, on any workflow where cache hits are expected often enough to matter. Anthropic describes the stored fingerprint as lightweight, scoped, and content-free, while the operational upside is large: the next miss stops being a blind bill spike and becomes one named divergence. That is especially useful on agent systems, where tool schemas, system prompts, and history replay are changing under active development.

I would still treat diagnostics as an operator feature rather than model logic. The model does not need to see or reason about these results. Your application does. Keep the response id in application state, watch for unavailable, and stabilise beta headers before you chase subtler causes. If you are also using Claude compaction, keep the same rule in mind: append cleanly, preserve byte stability, and do not rewrite old structure unless you are willing to lose the cache. The engineering lesson here is simple. Claude now tells you the first place the prefix moved. Teams that still debug cache misses by eyeballing prompts are choosing pain they no longer need.

Do this

Debug Claude cache misses without guessing

The aim is to turn a vague drop in cache_read_input_tokens into one concrete edit to the request shape.

  1. Enable diagnostics on every turn you expect to reuse

    Anthropic stores a fingerprint only when the diagnostics object is present. If you skip one turn, the next comparison can only return previous_message_not_found.

  2. Persist the response id with the conversation state

    The next request needs diagnostics.previous_message_id set to the prior diagnosed response id. Persist it in conversation metadata. The model should never carry it.

  3. Hold prompt-affecting parameters constant first

    Keep model, tool_choice, thinking, context_management, output_config, output_format, and active beta headers fixed while you debug. Anthropic groups several of these under unavailable, which is still enough to tell you the request envelope changed.

  4. Make system text and tools byte stable

    Move timestamps, request ids, and other per-turn values below the cache breakpoint. Send tools in a fixed order and serialize every input schema deterministically.

  5. Treat messages as append only

    Echo assistant content and tool_result blocks back verbatim. Rewriting, truncating, or reordering earlier history is how messages_changed appears.

  6. Read diagnostics before you read the bill

    Fix the first divergence the API names, then check usage.cache_read_input_tokens and cache_creation_input_tokens. Diagnostics tells you what changed. Usage tells you whether caching actually happened.

  7. Check prompt length if nothing diverged

    Anthropic's prompt caching page says short prompts are processed without caching and without an error. A structurally stable request can still miss if it never crossed the model's minimum cacheable length.

Frequently asked questions

Does Claude cache diagnostics store my prompt text?
No. Anthropic's cache diagnostics page says fingerprints contain only hashes and token-count estimates, never raw prompt content, and are scoped to your organization and workspace.
Does previous_message_not_found mean my prompt drifted?
No. Anthropic says it means no stored fingerprint exists for that previous_message_id. The common causes are that the earlier turn did not include the diagnostics object, it came from a different workspace, or too much time passed.
Does diagnostics null mean Claude missed the cache?
No. The cache diagnostics page says null can mean the request omitted diagnostics, previous_message_id was null on the first compared turn, or a comparison ran and found no divergence. Use usage.cache_read_input_tokens to confirm a real cache hit.
What does cache_miss_reason null mean inside the diagnostics object?
Anthropic says that shape means the comparison was still running when the response was serialized. It is inconclusive rather than a miss reason, so the next turn is the safer place to judge.
Why can usage.cache_read_input_tokens be zero even when diagnostics finds no divergence?
Prompt caching still has model-specific minimum lengths and explicit cache breakpoints. Anthropic's prompt caching page says shorter prompts are processed without caching and no error is returned.

Sources

  1. Cache diagnosticsAnthropic · 2026-09-30
  2. Prompt cachingAnthropic · 2026-09-30
  3. Claude release notes overviewAnthropic · 2026-09-23

claude-apiprompt-cachingdiagnosticsobservabilityagents