Gemini API: All about tokens
View original →Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.
Gemini API: All about tokens
<a target="_blank" href="https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Counting_Tokens.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" height=30/></a>
An understanding of tokens is central to using the Gemini API. This guide will provide a interactive introduction to what tokens are and how they are used in the Gemini API.
About tokens
LLMs break up their input and produce their output at a granularity that is smaller than a word, but larger than a single character or code-point.
These tokens can be single characters, like z, or whole words, like the. Long words may be broken up into several tokens. The set of all tokens used by the model is called the vocabulary, and the process of breaking down text into tokens is called tokenization.
For Gemini models, a token is equivalent to about 4 characters. 100 tokens are about 60-80 English words.
When billing is enabled, the price of a paid request is controlled by the number of input and output tokens, so knowing how to count your tokens is important.
Note: This notebook uses the Interactions API, the latest way to interact with Gemini models. Looking for the
generateContentversion? Check the archive branch.
Setup
Install SDK
Install the SDK from PyPI.
%pip install -U -q "google-genai>=2.9.0" # 2.0 for Interactions API
Setup your API key
To run the following cell, your API key must be stored in a Colab Secret named GEMINI_API_KEY. If you don't already have an API key, or you're not sure how to create a Colab Secret, see Authentication for a walkthrough.
from google.colab import userdata
GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')
Initialize SDK client
With the new SDK you now only need to initialize a client with your API key (or OAuth if using Vertex AI). The model is now set in each call.
from google import genai
client = genai.Client(api_key=GEMINI_API_KEY)
Tokens in the Gemini API
Context windows
The models available through the Gemini API have context windows that are measured in tokens. These define how much input you can provide, and how much output the model can generate, and combined are referred to as the "context window". This information is available directly through the API and in the models documentation.
In this example you can see the gemini-3.7-flash model has an 1M tokens context window. If you need more, Pro models have an even bigger 2M tokens context window.
MODEL_ID = "gemini-3.8-flash"
model_info = client.models.get(model=MODEL_ID)
print("Context window:",model_info.input_token_limit, "tokens")
print("Max output window:",model_info.output_token_limit, "tokens")
Counting tokens
The API provides an endpoint for counting the number of tokens in a request: client.models.count_tokens. You pass the same arguments as you would to client.models.interactions.create and the service will return the number of tokens in that request.
Select the model you want to use in this guide:
MODEL_ID = "gemini-3.8-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input": true, "isTemplate": true}
Text tokens
response = client.models.count_tokens(
model=MODEL_ID,
contents="What's the highest mountain in Africa?",
)
print("Prompt tokens:",response.total_tokens)
When you call client.interactions.create the response object has a usage field that includes the token count information.
interaction = client.interactions.create(
model=MODEL_ID,
input="The quick brown fox jumps over the lazy dog."
)
print(interaction.steps[-1].content[0].text)
# Note: The exact field names may differ in Interactions API
# Check interaction.usage for token information
print(interaction.usage)
In case you are using context caching, the number of cached token will be indicated in response.usage_metadata.cached_content_token_count.
Multi-modal tokens
All input to the API is tokenized, including images or other non-text modalities.
Images are considered to be a fixed size, so they consume a fixed number of tokens, regardless of their display or file size.
Video and audio files are converted to tokens at a fixed per second rate.
The current rates and token sizes can be found on the documentation
Inline content
You can pass media directly by URL:
from google.genai import types
response = client.models.count_tokens(
model=MODEL_ID,
contents=[
types.Part(
file_data=types.FileData(
file_uri="https://storage.googleapis.com/generativeai-downloads/images/jetpack.jpg",
mime_type="image/jpeg",
)
)
],
)
print("Prompt with image tokens:", response.total_tokens)
You can try with different images and should always get the same number of tokens, that is independent of their display or file size. Note that an extra token seems to be added, representing the empty prompt.
Files API
The model sees identical tokens if you upload parts of the prompt through the files API instead:
from google.genai import types
response = client.models.count_tokens(
model=MODEL_ID,
contents=types.Part.from_uri(
file_uri="https://storage.googleapis.com/generativeai-downloads/images/jetpack.jpg",
mime_type="image/jpeg",
),
)
print("Prompt with image tokens:", response.total_tokens)
Audio and video are each converted to tokens at a fixed rate of tokens per minute.
import subprocess
from google.genai import types
url = "https://storage.googleapis.com/generativeai-downloads/data/State_of_the_Union_Address_30_January_1961.mp3"
# Compute duration programmatically
cmd = [
"ffprobe",
"-v",
"error",
"-show_entries",
"format=duration",
"-of",
"default=noprint_wrappers=1:nokey=1",
url,
]
duration = float(subprocess.check_output(cmd).decode().strip())
response = client.models.count_tokens(
model=MODEL_ID,
contents=types.Part.from_uri(
file_uri=url,
mime_type="audio/mp3",
),
)
print("Prompt with audio tokens:", response.total_tokens)
print("Tokens per second:", response.total_tokens / duration)
As you can see this corresponds to about 32 tokens per second of audio.
Chat, tools and caching
Chat, tools and caching are currently not supported by the unified SDK count_tokens method. This notebook will be updated when that will be the case.
In the meantime you can still check the token used after the call using the usage_metadata from the response. Check the context caching documentation for an example.
Further reading
For more on token counting, check out the documentation or the API reference:
countTokensREST API reference,count_tokensPython API reference,
Related Articles
The Token Efficiency Index: A Peer-Benchmarked Composite Indicator for AI Token Efficiency
As artificial intelligence (AI) adoption accelerates across tech giants, AI-native startups, and non-technical organizations alike, a deceptively simple question remains hard to answer: is that...
A*-Decoding: Token-Efficient Inference Scaling
Inference-time scaling has emerged as a powerful alternative to parameter scaling for improving language model performance on complex reasoning tasks. While existing methods have shown strong...
Token-Efficient RL for LLM Reasoning
We propose reinforcement learning (RL) strategies tailored for reasoning in large language models (LLMs) under strict memory and compute limits, with a particular focus on compatibility with LoRA...
SkillReducer: Optimizing LLM Agent Skills for Token Efficiency
LLM-based coding agents rely on \emph{skills}, pre-packaged instruction sets that extend agent capabilities, yet every token of skill content injected into the context window incurs both monetary...