Gemini API: Context Caching Quickstart
View original →Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
Gemini API: Context Caching Quickstart
<a target="_blank" href="https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Caching.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" height=30/></a>
This notebook introduces context caching with the Gemini API and provides examples of interacting with the Apollo 11 transcript using the Python SDK. For a more comprehensive look, check out the caching guide.
Install dependencies
%pip install -U -q "google-genai>=2.9.0" # 2.0 for Interactions API
Setup your API key
To run the following cell, your API key must be stored in a Colab Secret named GEMINI_API_KEY. If you don't already have an API key, or you're not sure how to create a Colab Secret, see Authentication Image: image for a walkthrough.
from google.colab import userdata
from google import genai
GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')
client = genai.Client(api_key=GEMINI_API_KEY)
Upload a file
A common pattern with the Gemini API is to ask a number of questions of the same document. Context caching is designed to assist with this case, and can be more efficient by avoiding the need to pass the same tokens through the model for each new request.
This example will be based on the transcript from the Apollo 11 mission.
Start by downloading that transcript.
!wget -q https://storage.googleapis.com/generativeai-downloads/data/a11.txt
!head a11.txt
Now upload the transcript using the File API.
document = client.files.upload(file="a11.txt")
Cache the prompt
Next create a CachedContent object specifying the prompt you want to use, including the file and other fields you wish to cache. In this example the system_instruction has been set, and the document was provided in the prompt.
Note that caches are model specific. You cannot use a cache made with a different model as their tokenization might be slightly different.
MODEL_ID = "gemini-3.8-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input": true, "isTemplate": true}
apollo_cache = client.caches.create(
model=MODEL_ID,
config={
'contents': [document],
'system_instruction': 'You are an expert at analyzing transcripts.',
},
)
apollo_cache
from IPython.display import Markdown
display(Markdown(f"As you can see in the output, you just cached **{apollo_cache.usage_metadata.total_token_count}** tokens."))
Manage the cache expiry
Once you have a CachedContent object, you can update the expiry time to keep it alive while you need it.
from google.genai import types
client.caches.update(
name=apollo_cache.name,
config=types.UpdateCachedContentConfig(ttl="7200s") # 2 hours in seconds
)
apollo_cache = client.caches.get(name=apollo_cache.name) # Get the updated cache
apollo_cache
Use the cache for generation
To use the cache for generation, pass the cache name in GenerateContentConfig(cached_content=apollo_cache.name) when calling client.models.generate_content().
response = client.models.generate_content(
model=MODEL_ID,
contents='Find a lighthearted moment from this transcript',
config=types.GenerateContentConfig(
cached_content=apollo_cache.name,
)
)
display(Markdown(response.text))
You can inspect token usage through usage_metadata. Note that the cached prompt tokens are included in prompt_token_count, but excluded from the total_token_count.
response.usage_metadata
display(Markdown(f"""
As you can see in the `usage_metadata`, the token usage is split between:
* {response.usage_metadata.cached_content_token_count} tokens for the cache,
* {response.usage_metadata.prompt_token_count} tokens for the input (including the cache, so {response.usage_metadata.prompt_token_count - response.usage_metadata.cached_content_token_count} for the actual prompt),
* {response.usage_metadata.thoughts_token_count} tokens for the thinking process,
* {response.usage_metadata.candidates_token_count} tokens for the output,
* {response.usage_metadata.total_token_count} tokens in total.
"""))
You can ask new questions of the model, and the cache is reused.
chat = client.chats.create(
model=MODEL_ID,
config={"cached_content": apollo_cache.name}
)
response = chat.send_message(message="Give me a quote from the most important part of the transcript.")
display(Markdown(response.text))
response = chat.send_message(
message="What was recounted after that?",
config={"cached_content": apollo_cache.name}
)
display(Markdown(response.text))
response.usage_metadata
display(Markdown(f"""
As you can see in the `usage_metadata`, the token usage is split between:
* {response.usage_metadata.cached_content_token_count} tokens for the cache,
* {response.usage_metadata.prompt_token_count} tokens for the input (including the cache, so {response.usage_metadata.prompt_token_count - response.usage_metadata.cached_content_token_count} for the actual prompt),
* {response.usage_metadata.thoughts_token_count} tokens for the thinking process,
* {response.usage_metadata.candidates_token_count} tokens for the output,
* {response.usage_metadata.total_token_count} tokens in total.
"""))
Since the cached tokens are cheaper than the normal ones, it means this prompt was much cheaper that if you had not used caching. Check the pricing here for the up-to-date discount on cached tokens.
Delete the cache
The cache has a small recurring storage cost (cf. pricing) so by default it is only saved for an hour. In this case you even set it up for a shorter amont of time (using "ttl") of 2h.
Still, if you don't need you cache anymore, it is good practice to delete it proactively.
print(apollo_cache.name)
client.caches.delete(name=apollo_cache.name)
Next Steps
Useful API references:
If you want to know more about the caching API, you can check the full API specifications and the caching documentation.
Continue your discovery of the Gemini API
Check the File API notebook to know more about that API. The vision capabilities of the Gemini API are a good reason to use the File API and the caching. The Gemini API also has configurable safety settings that you might have to customize when dealing with big files.
Related Articles
Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks
Recent advancements in Large Language Model (LLM) agents have enabled complex multi-turn agentic tasks requiring extensive tool calling, where conversations can span dozens of API calls with...
CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs
Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps save precious compute resources and speeds up...
Auditing Prompt Caching in Language Model APIs
Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than non-cached prompts. These timing differences introduce the risk of...
Prompt Cache: Modular Attention Reuse for Low-Latency Inference
We present Prompt Cache, an approach for accelerating inference for large language models (LLM) by reusing attention states across different LLM prompts. Many input prompts have overlapping text...