---
title: How do you debug Claude prompt cache misses?
url: https://deepthinkingai.org/claude-cache-diagnostics/
published: 2026-09-30
author: Shekhar Singh
topic: AI Engineering
tags: claude-api, prompt-caching, diagnostics, observability, agents
site: DeepThinking AI
---

# How do you debug Claude prompt cache misses?

**Summary:** Claude cache diagnostics compares the current request with the response `id` of the previous diagnosed turn and reports the first divergence across model, system, tools, or messages. Since 23 September 2026 the `diagnostics` field is always present, so reliable debugging means reading that field together with `usage.cache_read_input_tokens`.

## Key takeaways
- Claude cache diagnostics stores a fingerprint only for requests that include the diagnostics object, then compares the next request by previous response id.
- Claude reports only the earliest divergence, so a model or system change can hide later tool or message drift until you fix the first mismatch.
- Claude's previous_message_not_found result does not mean the prompt changed; it usually means no stored fingerprint was available to compare.
- Claude's unavailable result covers other prompt-affecting parameters such as tool_choice, thinking, context_management, output_config, output_format, beta headers, or a conversation beyond the comparison horizon.
- Anthropic made the response diagnostics field always present on 23 September 2026, which means parsers should distinguish null, pending comparison, and a concrete miss reason.

Claude's cache diagnostics page answers a real pain point that the prompt
caching page left open: a cache miss used to look like a single number falling
to zero. On 23 September 2026 Anthropic also changed the response shape so
`POST /v1/messages` always includes a `diagnostics` field, with `null` when you
did not ask for a comparison. That sounds cosmetic. It is actually the detail
that makes the feature safe to automate, because a parser can now distinguish
"no diagnostics requested" from "comparison still pending" from "a divergence
was found".

## What does Claude cache diagnostics actually compare?

Anthropic's docs describe cache diagnostics as a comparison between consecutive
requests, keyed by response id rather than by conversation id. When a request
includes the `diagnostics` object, Claude stores a lightweight fingerprint for
that turn. On the next request, you pass the earlier response `id` as
`diagnostics.previous_message_id`, and the API compares the new request against
the stored fingerprint.

The docs say the fingerprint carries hashes and token count estimates. It does
not store raw prompt text. It is scoped to your organization and workspace.

The important correction is simpler. This comparison explains request
structure. Claude checks whether the model, system prompt, tools, and earlier
message history still line up. Cache confirmation still lives in
`usage.cache_read_input_tokens`. If you are already using
[prompt caching economics](/prompt-caching-economics/) to decide whether
caching is worth the effort, cache diagnostics is the mechanism that tells you
where that effort is leaking.

**What Claude compares before naming a cache miss**

1. Model selection
   Cache fingerprints are per model.
2. System prompt bytes
   Dynamic values here force a miss.
3. Tools array and schemas
   Order and JSON stability matter.
--- First divergence here ---
4. Earlier message history
   History must stay append only.
5. Later mismatches stay hidden
   Fix the first reason, then rerun.

Anthropic returns the earliest divergence only. That makes the result actionable, but it also means a model or system mismatch can hide tool or message drift until the earlier issue is fixed.

## What does each Claude cache miss reason mean?

Anthropic reports only the earliest divergence, which is exactly the behavior
you want during triage. `model_changed` means the model switched between turns,
often because a router, A/B test, or fallback path picked another model.
`system_changed` means the `system` parameter was byte different, which the docs
call out as a common side effect of interpolating timestamps or request ids.
`tools_changed` covers added, removed, or reordered tools and
non-deterministically serialized `input_schema` JSON. `messages_changed` means
the model, system, and tools still match, but earlier history was edited,
reordered, or truncated instead of appended to.

Two results matter because they are easy to misread. `previous_message_not_found`
means no stored fingerprint existed for the supplied response id. `unavailable`
means Claude could not produce a comparison even though the request still ran.
Anthropic names concrete causes here: other prompt-affecting parameters changed,
including `tool_choice`, `thinking`, `context_management`, `output_config`,
`output_format`, or the set of active beta headers, or the conversation drifted
beyond the comparison horizon. The fix is to stabilise the request envelope,
then rerun. A later problem may still be waiting behind the first one.

**How a diagnosed Claude turn becomes the next comparison**

```mermaid
sequenceDiagram
    participant Client as Client
    participant ClaudeAPI as Claude API
    Client->>ClaudeAPI: Send request with diagnostics enabled
    ClaudeAPI->>Client: Return response id
    Client->>ClaudeAPI: Send next request with previous_message_id
    ClaudeAPI->>Client: Return diagnostics plus usage counters
```

- Send request with diagnostics enabled: API stores a fingerprint for this turn.
- Return response id: Save it beside the conversation.
- Send next request with previous_message_id: Same model and stable prefix expected.
- Return diagnostics plus usage counters: Compare structure, then confirm the hit.

Cache diagnostics is a two-turn mechanism. The first diagnosed turn creates the fingerprint. The next turn supplies that response id so Anthropic can compare the new request against the stored structure.

## Why can Claude say previous_message_not_found when nothing changed?

Because that result reflects missing comparison state. It says nothing yet about
prompt drift.
Anthropic says plainly that no fingerprint exists when the earlier request did
not include the `diagnostics` object, came from another workspace, or is too
old to retain. In other words, `previous_message_not_found` is often a logging
or retention issue that happens before any structural comparison could begin.

That is why I would not treat this result as evidence of a cache regression in
itself. The safer reading is that the debugging chain is broken. Re-enable
diagnostics on every turn, keep consecutive turns close together, and make sure
your application persists the previous response id beside the conversation. The
same discipline matters for other cache-preserving features on Claude. A
[mid-conversation tool change](/claude-mid-conversation-tool-changes/) only
helps if the application keeps the surrounding request stable enough for the
comparison to mean anything. The id plumbing is dull, but it is the part that
turns the API from "I think we changed something" into a concrete explanation.

## Why should you read Claude diagnostics with usage counters?

Because a structural comparison and a real cache hit answer different questions.
Anthropic's docs say the `diagnostics` field can be `null` when you did not ask
for diagnostics, when `previous_message_id` was `null` on the first compared
turn, or when a comparison found no divergence. They also document a second
shape, `{"cache_miss_reason": null}`, for comparisons that were still running
when the response was serialized. Neither shape tells you by itself whether
cached tokens were actually read.

The usage counters do. `usage.cache_read_input_tokens` shows how many prompt
tokens came from cache, and `usage.cache_creation_input_tokens` shows how many
were written into a new cache entry. Anthropic's prompt caching page adds one
more operational trap: prompts below the model's minimum cacheable length are
processed without caching and without an error. That means you can have a stable
request and still see zero read tokens. Use diagnostics to find drift. Use
usage to confirm that the cached prefix existed, was long enough, and was
actually reused. That division is the key insight the release note summary does
not make explicit.

## Should you enable Claude cache diagnostics on every cached conversation?

My view is yes, on any workflow where cache hits are expected often enough to
matter. Anthropic describes the stored fingerprint as lightweight, scoped, and
content-free, while the operational upside is large: the next miss stops being
a blind bill spike and becomes one named divergence. That is especially useful
on agent systems, where tool schemas, system prompts, and history replay are
changing under active development.

I would still treat diagnostics as an operator feature rather than model logic.
The model does not need to see or reason about these results. Your application
does. Keep the response id in application state, watch for `unavailable`, and
stabilise beta headers before you chase subtler causes. If you are also using
[Claude compaction](/claude-compaction-modes/), keep the same rule in mind:
append cleanly, preserve byte stability, and do not rewrite old structure
unless you are willing to lose the cache. The engineering lesson here is simple.
Claude now tells you the first place the prefix moved. Teams that still debug
cache misses by eyeballing prompts are choosing pain they no longer need.

<ReadNext
  href="/prompt-caching-economics/"
  kicker="Go deeper"
  title="When does prompt caching actually save money?"
  note="Use the cache-miss diagnosis here with the cost model there, so you know both where the prefix drifted and whether fixing it changes the bill enough to matter."
/>

## Debug Claude cache misses without guessing

The aim is to turn a vague drop in cache_read_input_tokens into one concrete edit to the request shape.

1. **Enable diagnostics on every turn you expect to reuse**: Anthropic stores a fingerprint only when the diagnostics object is present. If you skip one turn, the next comparison can only return previous_message_not_found.
2. **Persist the response id with the conversation state**: The next request needs diagnostics.previous_message_id set to the prior diagnosed response id. Persist it in conversation metadata. The model should never carry it.
3. **Hold prompt-affecting parameters constant first**: Keep model, tool_choice, thinking, context_management, output_config, output_format, and active beta headers fixed while you debug. Anthropic groups several of these under unavailable, which is still enough to tell you the request envelope changed.
4. **Make system text and tools byte stable**: Move timestamps, request ids, and other per-turn values below the cache breakpoint. Send tools in a fixed order and serialize every input schema deterministically.
5. **Treat messages as append only**: Echo assistant content and tool_result blocks back verbatim. Rewriting, truncating, or reordering earlier history is how messages_changed appears.
6. **Read diagnostics before you read the bill**: Fix the first divergence the API names, then check usage.cache_read_input_tokens and cache_creation_input_tokens. Diagnostics tells you what changed. Usage tells you whether caching actually happened.
7. **Check prompt length if nothing diverged**: Anthropic's prompt caching page says short prompts are processed without caching and without an error. A structurally stable request can still miss if it never crossed the model's minimum cacheable length.


## Frequently asked questions

### Does Claude cache diagnostics store my prompt text?

No. Anthropic's cache diagnostics page says fingerprints contain only hashes and token-count estimates, never raw prompt content, and are scoped to your organization and workspace.

### Does previous_message_not_found mean my prompt drifted?

No. Anthropic says it means no stored fingerprint exists for that previous_message_id. The common causes are that the earlier turn did not include the diagnostics object, it came from a different workspace, or too much time passed.

### Does diagnostics null mean Claude missed the cache?

No. The cache diagnostics page says null can mean the request omitted diagnostics, previous_message_id was null on the first compared turn, or a comparison ran and found no divergence. Use usage.cache_read_input_tokens to confirm a real cache hit.

### What does cache_miss_reason null mean inside the diagnostics object?

Anthropic says that shape means the comparison was still running when the response was serialized. It is inconclusive rather than a miss reason, so the next turn is the safer place to judge.

### Why can usage.cache_read_input_tokens be zero even when diagnostics finds no divergence?

Prompt caching still has model-specific minimum lengths and explicit cache breakpoints. Anthropic's prompt caching page says shorter prompts are processed without caching and no error is returned.


## Sources
- [Cache diagnostics](https://platform.claude.com/docs/en/build-with-claude/cache-diagnostics). Anthropic, 2026-09-30
- [Prompt caching](https://platform.claude.com/docs/en/build-with-claude/prompt-caching). Anthropic, 2026-09-30
- [Claude release notes overview](https://platform.claude.com/docs/en/release-notes/overview). Anthropic, 2026-09-23

---
Canonical HTML: https://deepthinkingai.org/claude-cache-diagnostics/