@tokalator4.2.0…installs…downloads
Gemini CookbookGeneral

Use Gemini thinking

View original →
Copyright 2026 Google LLC.
#@title Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# https://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

Use Gemini thinking

Note: This notebook uses the Interactions API, the latest way to interact with Gemini models. Looking for the generateContent version? Check the archive branch.

<a target="_blank" href="https://colab.research.google.com/github/google-gemini/cookbook/blob/main/quickstarts/Get_started_thinking.ipynb"><img src="https://colab.research.google.com/assets/colab-badge.svg" height=30/></a>

All Gemini models are trained to do a thinking process (or reasoning) before getting to a final answer. As a result, those models usually get better results on harder tasks that require multiple processing steps: complex math, coding, reasoning over multi-step instructions and multimodal understanding.

While thinking is always on, you can configure the amount of thinking the model does by using thinking levels (minimal, low, medium, high). This lets you balance between response speed/cost and reasoning depth depending on your use case.

Understanding thinking models

Thinking models are optimized for complex tasks that need multiple rounds of strategizing and iteratively solving.

You can control the thinking effort using thinking levels:

Thinking LevelUse Case
High (default)Complex reasoning, math, coding, multi-step problems
MediumGood balance between speed and reasoning
LowFaster responses with some reasoning
MinimalFastest responses, minimal reasoning (roughly equivalent to "off")

You can set the thinking level using types.ThinkingConfig(thinking_level=types.ThinkingLevel.HIGH).

Setup

This section install the SDK, set it up using your API key, imports the relevant libs, downloads the sample videos and upload them to Gemini.

Just collapse (click on the little arrow on the left of the title) and run this section if you want to jump straight to the examples (just don't forget to run it otherwise nothing will work).

Install SDK

The Google Gen AI SDK provides programmatic access to Gemini models using both the Google AI for Developers and Vertex AI APIs. With a few exceptions, code that runs on one platform will run on both. This means that you can prototype an application using the Developer API and then migrate the application to Vertex AI without rewriting your code.

More details about this new SDK on the documentation or in the Getting started Image: image notebook.

%pip install -U -q "google-genai>=2.9.0" # 2.0 is needed to use the interactions API

Setup your API key

To run the following cell, your API key must be stored in a Colab Secret named GEMINI_API_KEY. If you don't already have an API key, or you're not sure how to create a Colab Secret, see Authentication Image: image for a walkthrough.

from google.colab import userdata

GEMINI_API_KEY=userdata.get('GEMINI_API_KEY')

Initialize SDK client

With the new SDK you now only need to initialize a client with your API key (or OAuth if using Vertex AI). The model is now set in each call.

from google import genai
from google.genai import types

client = genai.Client(api_key=GEMINI_API_KEY)
MODEL_ID = "gemini-3.8-flash" # @param ["gemini-3.1-pro-preview", "gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash-lite", "gemini-2.5-pro"] {"allow-input": true, "isTemplate": true}

Imports

import base64
import json
from IPython.display import display, HTML, Markdown
from PIL import Image

Using thinking models

Here are some complex examples of what Gemini thinking models can solve.

In each of them you can select different models to see how they compare.

In some cases, you'll still get a good answer even with lower thinking levels. In that case, re-run the example with minimal thinking to see how much the model benefits from deeper reasoning.

Using thinking levels

You can start by asking the model to explain a concept and see how it does reasoning before answering.

By default, the model uses the high thinking level, which lets it dynamically adjust its reasoning depth based on the complexity of the request.

prompt = """
    You are playing the 20 question game. You know that what you are looking for
    is a aquatic mammal that doesn't live in the sea, is venomous and that's
    smaller than a cat. What could that be and how could you make sure?
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
)

display(Markdown(interaction.output_text))

Looking to the response metadata, you can see not only the amount of tokens on your input and the amount of tokens used for the response, but also the amount of tokens used for the thinking step - As you can see here, the model used around 1400 tokens in the thinking steps:

print("Prompt tokens:", interaction.usage.total_input_tokens)
print("Thoughts tokens:", interaction.usage.total_thought_tokens)
print("Output tokens:", interaction.usage.total_output_tokens)
print("Total tokens:", interaction.usage.total_tokens)

Low thinking

You can set thinking to low to get faster responses with reduced reasoning. You'll see that in this case, the model spends less time thinking through the problem.

if "-pro" not in MODEL_ID:
  prompt = """
      You are playing the 20 question game. You know that what you are looking for
      is a aquatic mammal that doesn't live in the sea, is venomous and that's
      smaller than a cat. What could that be and how could you make sure?
  """

  interaction = client.interactions.create(
      model=MODEL_ID,
      input=prompt,
      generation_config={
          "thinking_level": "low",
      },
  )

  display(Markdown(interaction.output_text))

else:
  print("You can't disable thinking for pro models.")

Now you can see that the response is faster as the model didn't perform any thinking step. Also you can see that no tokens were used for the thinking step:

print("Prompt tokens:", interaction.usage.total_input_tokens)
print("Thoughts tokens:", interaction.usage.total_thought_tokens)
print("Output tokens:", interaction.usage.total_output_tokens)
print("Total tokens:", interaction.usage.total_tokens)

Now, try with a complex physics problem. First with low thinking to see how the model performs:

if "-pro" not in MODEL_ID:
  prompt = """
      A cantilever beam of length L=3m has a rectangular cross-section
      (width b=0.1m, height h=0.2m) and is made of steel (E=200 GPa).
      It is subjected to a uniformly distributed load w=5 kN/m along its entire
      length and a point load P=10 kN at its free end.
      Calculate the maximum bending stress (σ_max).
  """

  interaction = client.interactions.create(
      model=MODEL_ID,
      input=prompt,
      generation_config={
          "thinking_level": "low",
      },
  )

  display(Markdown(interaction.output_text))

else:
  print("You can't disable thinking for pro models.")

You can see that with low thinking, the model uses fewer tokens for reasoning:

print("Prompt tokens:", interaction.usage.total_input_tokens)
print("Thoughts tokens:", interaction.usage.total_thought_tokens)
print("Output tokens:", interaction.usage.total_output_tokens)
print("Total tokens:", interaction.usage.total_tokens)

Then you can use high thinking to see how the model performs with full reasoning.

You can see that, even for the same prompt, the depth and consistency of the answer improves significantly when the model is allowed to think more deeply.

NOTE: You can always see how many tokens were used for the thinking step in the usage_metadata.

prompt = """
    A cantilever beam of length L=3m has a rectangular cross-section
    (width b=0.1m, height h=0.2m) and is made of steel (E=200 GPa).
    It is subjected to a uniformly distributed load w=5 kN/m along its entire
    length and a point load P=10 kN at its free end.
    Calculate the maximum bending stress (σ_max).
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    generation_config={
        "thinking_level": "high",
    },
)

display(Markdown(interaction.output_text))

Now you can see that the model used significantly more tokens for reasoning:

print("Prompt tokens:", interaction.usage.total_input_tokens)
print("Thoughts tokens:", interaction.usage.total_thought_tokens)
print("Output tokens:", interaction.usage.total_output_tokens)
print("Total tokens:", interaction.usage.total_tokens)

Keep in mind that higher thinking levels mean the model will spend more time reasoning, which means longer computation time and a more expensive request.

<a name="geometry"></a>

Working with multimodal problems

This geometry problem requires complex reasoning and is also using Gemini multimodal abilities to read the image.

!wget https://storage.googleapis.com/generativeai-downloads/images/geometry.png -O geometry.png -q

geometry_image = Image.open("geometry.png").resize((256,256))
geometry_image
prompt = "What's the area of the overlapping region?"

with open("geometry.png", "rb") as f:
  geometry_b64 = base64.b64encode(f.read()).decode("utf-8")

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "image", "data": geometry_b64, "mime_type": "image/png"},
        {"type": "text", "text": prompt},
    ],
    generation_config={
        "thinking_level": "high",
    },
)

display(Markdown(interaction.output_text))

Solving brain teasers

Here's another brain teaser based on an image, this time it looks like a mathematical problem, but it cannot actually be solved mathematically. If you check the thoughts of the model you'll see that it will realize it and come up with an out-of-the-box solution.

!wget https://storage.googleapis.com/generativeai-downloads/images/pool.png -O pool.png -q

pool_image = Image.open("pool.png").resize((256,256))
pool_image

First you can check how the model performs with low thinking:

with open("pool.png", "rb") as f:
  pool_b64 = base64.b64encode(f.read()).decode("utf-8")

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "image", "data": pool_b64, "mime_type": "image/png"},
        {
            "type": "text",
            "text": "How do I use those three pool balls to sum up to 30?",
        },
    ],
    generation_config={
        "thinking_level": "low",
    },
)

display(Markdown(interaction.output_text))

As you can notice, the model struggled to find a way to get to the result — and ended up suggesting to use different pool balls.

Now you can use high thinking to solve the riddle:

prompt = "How do I use those three pool balls to sum up to 30?"

with open("pool.png", "rb") as f:
  pool_b64 = base64.b64encode(f.read()).decode("utf-8")

interaction = client.interactions.create(
    model=MODEL_ID,
    input=[
        {"type": "image", "data": pool_b64, "mime_type": "image/png"},
        {"type": "text", "text": prompt},
    ],
    generation_config={
        "thinking_level": "high",
    },
)

display(Markdown(interaction.output_text))

<a name="math"></a>

Solving a math puzzle with high thinking

This is typically a case where you want to use high thinking, as the model needs to explore many directions before finding the right answer.

prompt = """
   How can you obtain 565 with 10 8 3 7 1 and 5 and the common operations?
   You can only use a number once.
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    generation_config={
        "thinking_level": "high",
    },
)

display(Markdown(interaction.output_text))

<a name="thoughts_summaries"></a>

Working thoughts summaries

Summaries of the model's thinking reveal its internal problem-solving pathway. Users can leverage this feature to check the model's strategy and remain informed during complex tasks.

For more details about Gemini thinking capabilities, take a look at the Gemini models thinking guide.

prompt = """
  Alice, Bob, and Carol each live in a different house on the same street: red, green, and blue.
  The person who lives in the red house owns a cat.
  Bob does not live in the green house.
  Carol owns a dog.
  The green house is to the left of the red house.
  Alice does not own a cat.
  Who lives in each house, and what pet do they own?
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    generation_config={
        "thinking_level": "high",
        "thinking_summaries": "auto",
    },
)

You can check both the thought summaries and the final model response:

for step in interaction.steps:
  if step.type == "thought" and getattr(step, "summary", None):
    summary_items = (
        step.summary
        if isinstance(step.summary, list)
        else [step.summary]
    )
    for item in summary_items:
      text = getattr(item, "text", str(item))
      display(Markdown(f"## **Thoughts summary:**\n\n{text}"))
      print()
  elif step.type == "model_output" and step.content:
    for content in step.content:
      if getattr(content, "thought", False):
        display(Markdown("## **Thoughts summary:**"))
        display(Markdown(content.text))
        print()
      elif getattr(content, "text", None):
        display(Markdown("## **Answer:**"))
        display(Markdown(content.text))

You can also use see the thought summaries in streaming experiences:

prompt = """
    Alice, Bob, and Carol each live in a different house on the same street: red,
    green, and blue.
    The person who lives in the red house owns a cat.
    Bob does not live in the green house.
    Carol owns a dog.
    The green house is to the left of the red house.
    Alice does not own a cat.
    Who lives in each house, and what pet do they own?
"""

thoughts = ""
answer = ""

for event in client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    generation_config={
        "thinking_level": "high",
        "thinking_summaries": "auto",
    },
    stream=True,
):
  if hasattr(event, "delta") and event.delta:
    if getattr(event.delta, "type", None) == "thought" or getattr(
        event.delta, "thought", False
    ):
      thoughts += getattr(event.delta, "text", "") or ""
    elif getattr(event.delta, "text", None):
      answer += event.delta.text

if thoughts:
  display(Markdown("## **Thoughts summary:**"))
  display(Markdown(thoughts))
  print()
display(Markdown("## **Answer:**"))
display(Markdown(answer))

Working with Gemini thinking models and tools

Gemini thinking models are compatible with the tools and capabilities inherent to the Gemini ecosystem. This compatibility allows them to interface with external environments, execute computational code, or retrieve real-time data, subsequently incorporating such information into their analytical framework and concluding statements.

<a name="code_execution"></a>

Solving a problem using the code execution tool

This example shows how to use the code execution tool to solve a problem. The model will generate the code and then execute it to get the final answer.

prompt = """
    What are the best ways to sort a list of n numbers from 0 to m?
    Generate and run Python code for three different sort algorithms.
    Provide the final comparison between algorithm clearly.
    Is one of them linear?
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    tools=[{"type": "code_execution"}],
    generation_config={
        "thinking_level": "high",
    },
)

Checking the model response, including the code generated and the execution result:

from IPython.display import display, HTML, Markdown

for step in interaction.steps:
  if step.type == "model_output" and step.content:
    for content in step.content:
      if getattr(content, "text", None) is not None:
        display(Markdown(content.text))
  elif step.type == "code_execution_call":
    arguments = getattr(step, "arguments", None)
    code = dict(arguments).get("code", "") if arguments else ""
    display(
        HTML(
            '<pre style="background-color: #1a1a2e; color: #16c60c;'
            f' padding: 10px;">{code}</pre>'
        )
    )
  elif step.type == "code_execution_result":
    if hasattr(step, "result") and step.result:
      display(
          HTML(
              '<pre style="background-color: #2d2d44; color: white;'
              f' padding: 10px;">{step.result}</pre>'
          )
      )

<a name="google_search"></a>

Thinking with search tool

Search grounding is a great way to improve the quality of the model responses by giving it the ability to search for the latest information using Google Search. Check the dedicated guide for more details on that feature.

prompt = """
    What were the major scientific breakthroughs announced last month? Use your
    critical thinking and only list what's really incredible and not just an
    overinfluated title.
"""

interaction = client.interactions.create(
    model=MODEL_ID,
    input=prompt,
    tools=[{"type": "google_search"}],
    generation_config={
        "thinking_level": "high",
        "thinking_summaries": "auto",
    },
)

Then you can check all information:

  • the model thoughts summary
  • the model answer
  • and the Google Search reference
for step in interaction.steps:
  if step.type == "thought" and getattr(step, "summary", None):
    summary_items = (
        step.summary
        if isinstance(step.summary, list)
        else [step.summary]
    )
    for item in summary_items:
      text = getattr(item, "text", str(item))
      display(Markdown(f"## **Thoughts summary:**\n\n{text}"))
      print()
  elif step.type == "model_output" and step.content:
    for content in step.content:
      if getattr(content, "thought", False):
        display(Markdown("## **Thoughts summary:**"))
        display(Markdown(content.text))
        print()
      elif getattr(content, "text", None):
        display(Markdown("## **Answer:**"))
        display(Markdown(content.text))

<a name="thinking_level"></a>

Thinking levels reference

Thinking levels provide a simple way to control the amount of reasoning:

# Set thinking level
config = types.GenerateContentConfig(
    thinking_config=types.ThinkingConfig(
        thinking_level=types.ThinkingLevel.HIGH  # MINIMAL, LOW, MEDIUM, HIGH
    )
)
# @title Run this cell to set everything up if you jump directly to this section
from google import genai
from google.colab import userdata
from google.genai import types
from IPython.display import display, HTML, Markdown

client = genai.Client(api_key=userdata.get("GEMINI_API_KEY"))

# Select the Gemini 3 model

GEMINI_3_MODEL_ID = "gemini-3.8-flash"  # @param ["gemini-3.1-pro-preview", "gemini-3.8-flash", "gemini-3.7-flash", "gemini-3.6-flash", "gemini-3.5-flash-lite"] {"allow-input": true, "isTemplate": true}

You can set the thinking level to low, medium or high (default). This will indicate to the model how much thinking it is allowed to do. Since the thinking process stays dynamic, high doesn't necessarily mean the model will think a lot — it just means it's allowed to think as much as it needs.

prompt = """
  Find what I'm thinking of:
    It moves, but doesn't walk, run, or swim.
    It has no fixed shape and if cut into pieces, those pieces will keep living and moving.
    It has no brain but can solve complex mazes.
"""

# Thinking levels can be "low", "medium", or "high"
thinking_level = "high"  # @param ["low", "medium", "high"]

interaction = client.interactions.create(
    model=GEMINI_3_MODEL_ID,
    input=prompt,
    generation_config={
        "thinking_level": thinking_level,
        "thinking_summaries": "auto",
    },
)

for step in interaction.steps:
  if step.type == "thought" and getattr(step, "summary", None):
    summary_items = (
        step.summary
        if isinstance(step.summary, list)
        else [step.summary]
    )
    for item in summary_items:
      text = getattr(item, "text", str(item))
      display(Markdown("### Thought summary:"))
      display(Markdown(text))
      print()
  elif step.type == "model_output" and step.content:
    for content in step.content:
      if getattr(content, "text", None):
        display(Markdown("### Answer:"))
        display(Markdown(content.text))
        print()

if hasattr(interaction, "usage") and interaction.usage:
  print(
      f"We used {interaction.usage.total_thought_tokens} tokens for the"
      f" thinking phase and {interaction.usage.total_input_tokens} for"
      " the prompt."
  )

<a name="gemini3migration"></a>

Legacy: Migrating from thinking_budget to thinking_level

If you were previously using thinking_budget (a feature from generateContent), here's how to migrate:

  • If you were using thinking_budget=0 → use ThinkingLevel.MINIMAL
  • If you were using a low thinking_budget (e.g., 1024) → use ThinkingLevel.LOW
  • If you were using a medium thinking_budget (e.g., 4096-8192) → use ThinkingLevel.MEDIUM
  • If you were using the maximum thinking_budget or adaptive (default) → use ThinkingLevel.HIGH
  • If you were using dynamic thinking → use ThinkingLevel.HIGH (the model will still adjust dynamically)

Note: thinking_budget is a legacy feature from the generateContent API that can still be used to fine-tune thinking behavior. See the generateContent notebook for details.

Next Steps

Try Gemini 3 and other Gemini models in Google AI Studio, and learn more about Prompting for thinking models.

For more examples of the Gemini capabilities, check the other Cookbook examples. You'll learn how to use the Live API Image: image, juggle with multiple tools Image: image or use Gemini spatial understanding Image: image abilities.

tokenspromptsstreaming

Related Articles