DictionaryToken Economics See it in actionContext budget tool
Max Output Tokens
A cap on how long the AI's reply can be; if it hits the cap, the answer simply stops.
Definition
The maximum number of tokens a model may generate in one response, set per request with max_tokens and capped per model. In the site's catalog, current Claude and GPT models allow 128K output tokens, Gemini models 65,536, and Claude Haiku 4.5 and Opus 4.5 64K. Thinking counts toward the limit, and a response that reaches it stops with stop_reason max_tokens, possibly mid-sentence.
Example
With max_tokens: 1024, a long explanation is cut off mid-sentence and the response reports stop_reason: "max_tokens".