DictionaryContext Management See it in actionContext budget tool
Context Window
The amount of text an AI can keep in mind at once, like the size of a desk where everything must fit.
Definition
The maximum number of tokens (input plus output) a model can handle in one request: everything the model can see while generating. In the site's catalog most current models offer about 1M tokens (Claude Opus 5.5 and Sonnet 5.5: 1M; Gemini 3.8 Flash: 1,048,576; GPT-6 and GPT-5.6 models: 1.05M), while older models such as Claude Haiku 4.5 and Opus 4.5 have 200K. On the current Claude tokenizer, 1M tokens is roughly 555K words.
Example
A 1M-token window holds roughly 555K words. Paste a large codebase plus a long chat history and older content starts crowding out new work.
Appears in
- Diagnosing and Mitigating Context Rot in Long-horizon SearcharXiv
- SkillReducer: Optimizing LLM Agent Skills for Token EfficiencyarXiv
- The Root Theorem of Context EngineeringarXiv
- The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-CalculusarXiv
- What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and MitigationsarXiv
- Discrete Prompt Compression with Reinforcement LearningarXiv