@tokalator4.2.0…installs…downloads
Gemini CookbookToken Optimization

Gemini API: All about tokens

View original →
Copyright 2026 Google LLC.
# @title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

Gemini API: All about tokens

<a target="_blank" href="https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Counting_Tokens.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" height=30/></a>

An understanding of tokens is central to using the Gemini API. This guide will provide a interactive introduction to what tokens are and how they are used in the Gemini API.

About tokens

LLMs break up their input and produce their output at a granularity that is smaller than a word, but larger than a single character or code-point.

These tokens can be single characters, like z, or whole words, like the. Long words may be broken up into several tokens. The set of all tokens used by the model is called the vocabulary, and the process of breaking down text into tokens is called tokenization.

For Gemini models, a token is equivalent to about 4 characters. 100 tokens are about 60-80 English words.

When billing is enabled, the price of a paid request is controlled by the number of input and output tokens, so knowing how to count your tokens is important.

Note: This notebook uses the Interactions API, the latest way to interact with Gemini models. Looking for the generateContent version? Check the archive branch.

Setup

Install SDK

Install the SDK from PyPI.

%pip install -U -q "google-genai>=2.9.0"  # 2.0 for Interactions API

Setup your API key

To run the following cell, your API key must be stored in a Colab Secret named GEMINI_API_KEY. If you don't already have an API key, or you're not sure how to create a Colab Secret, see Authentication for a walkthrough.

from google.colab import userdata

GEMINI_API_KEY = userdata.get('GEMINI_API_KEY')

Initialize SDK client

With the new SDK you now only need to initialize a client with your API key (or OAuth if using Vertex AI). The model is now set in each call.

from google import genai

client = genai.Client(api_key=GEMINI_API_KEY)

Tokens in the Gemini API

Context windows

The models available through the Gemini API have context windows that are measured in tokens. These define how much input you can provide, and how much output the model can generate, and combined are referred to as the "context window". This information is available directly through the API and in the models documentation.

In this example you can see the gemini-3.7-flash model has an 1M tokens context window. If you need more, Pro models have an even bigger 2M tokens context window.

MODEL_ID = "gemini-3.8-flash"

model_info = client.models.get(model=MODEL_ID)

print("Context window:",model_info.input_token_limit, "tokens")
print("Max output window:",model_info.output_token_limit, "tokens")

Counting tokens

The API provides an endpoint for counting the number of tokens in a request: client.models.count_tokens. You pass the same arguments as you would to client.models.interactions.create and the service will return the number of tokens in that request.

Select the model you want to use in this guide:

MODEL_ID = "gemini-3.8-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input": true, "isTemplate": true}

Text tokens

response = client.models.count_tokens(
    model=MODEL_ID,
    contents="What's the highest mountain in Africa?",
)
print("Prompt tokens:",response.total_tokens)

When you call client.interactions.create the response object has a usage field that includes the token count information.

interaction = client.interactions.create(
    model=MODEL_ID,
    input="The quick brown fox jumps over the lazy dog."
)
print(interaction.steps[-1].content[0].text)
# Note: The exact field names may differ in Interactions API
# Check interaction.usage for token information
print(interaction.usage)

In case you are using context caching, the number of cached token will be indicated in response.usage_metadata.cached_content_token_count.

Multi-modal tokens

All input to the API is tokenized, including images or other non-text modalities.

Images are considered to be a fixed size, so they consume a fixed number of tokens, regardless of their display or file size.

Video and audio files are converted to tokens at a fixed per second rate.

The current rates and token sizes can be found on the documentation

Inline content

You can pass media directly by URL:

from google.genai import types

response = client.models.count_tokens(
    model=MODEL_ID,
    contents=[
        types.Part(
            file_data=types.FileData(
                file_uri="https://storage.googleapis.com/generativeai-downloads/images/jetpack.jpg",
                mime_type="image/jpeg",
            )
        )
    ],
)

print("Prompt with image tokens:", response.total_tokens)

You can try with different images and should always get the same number of tokens, that is independent of their display or file size. Note that an extra token seems to be added, representing the empty prompt.

Files API

The model sees identical tokens if you upload parts of the prompt through the files API instead:

from google.genai import types

response = client.models.count_tokens(
    model=MODEL_ID,
    contents=types.Part.from_uri(
        file_uri="https://storage.googleapis.com/generativeai-downloads/images/jetpack.jpg",
        mime_type="image/jpeg",
    ),
)

print("Prompt with image tokens:", response.total_tokens)

Audio and video are each converted to tokens at a fixed rate of tokens per minute.

import subprocess
from google.genai import types

url = "https://storage.googleapis.com/generativeai-downloads/data/State_of_the_Union_Address_30_January_1961.mp3"

# Compute duration programmatically
cmd = [
    "ffprobe",
    "-v",
    "error",
    "-show_entries",
    "format=duration",
    "-of",
    "default=noprint_wrappers=1:nokey=1",
    url,
]
duration = float(subprocess.check_output(cmd).decode().strip())

response = client.models.count_tokens(
    model=MODEL_ID,
    contents=types.Part.from_uri(
        file_uri=url,
        mime_type="audio/mp3",
    ),
)

print("Prompt with audio tokens:", response.total_tokens)
print("Tokens per second:", response.total_tokens / duration)

As you can see this corresponds to about 32 tokens per second of audio.

Chat, tools and caching

Chat, tools and caching are currently not supported by the unified SDK count_tokens method. This notebook will be updated when that will be the case.

In the meantime you can still check the token used after the call using the usage_metadata from the response. Check the context caching documentation for an example.

Further reading

For more on token counting, check out the documentation or the API reference:

tokenspromptscaching

Related Articles