@tokalator4.2.0…installs…downloads
Dictionary
Caching

Prompt Caching

The AI service remembers the unchanged start of your request so repeat requests are cheaper and faster, like a saved draft.

Definition

A technique where frequently-used prompt prefixes are stored server-side, allowing subsequent requests with the same prefix to be processed at reduced cost. Cache reads cost 90% or more below base input on current models. Requires a minimum token threshold to activate.

Example

Mark a 20K-token system prompt with "cache_control": {"type": "ephemeral"}; calls within 5 minutes read it at a fraction of the input price.

Appears in

See it in actionPrompt caching tool