DictionaryCaching See it in actionPrompt caching tool
Prompt Caching
The AI service remembers the unchanged start of your request so repeat requests are cheaper and faster, like a saved draft.
Definition
A technique where frequently-used prompt prefixes are stored server-side, allowing subsequent requests with the same prefix to be processed at reduced cost. Cache reads cost 90% or more below base input on current models. Requires a minimum token threshold to activate.
Example
Mark a 20K-token system prompt with "cache_control": {"type": "ephemeral"}; calls within 5 minutes read it at a fraction of the input price.
Appears in
- Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic TasksarXiv
- Auditing Prompt Caching in Language Model APIsarXiv
- Efficient Prompt Caching via Embedding SimilarityarXiv
- Prompt caching with the Claude APIClaude Cookbook
- Speculative Prompt CachingClaude Cookbook
- CacheProbe: Auditing Prompt Cache Isolation in Gateway APIsarXiv