DeepThinking AI

Tag

production

Covered in AI Engineering, where the background and the sources for this subject live.

AI Engineering

When does prompt caching actually save money?

Prompt caching stores a prefix of your prompt so later requests reuse it instead of reprocessing it. Reads are far cheaper than base input tokens, but writing the cache costs a premium and entries expire. It pays whenever a large stable prefix is reused several times inside the TTL, and loses on one-shot traffic.

5 min read