Learn token economics
Short lessons on tokens, context windows and cost, with runnable code.
What Is a Prompt?
Instructions made of tokens
A prompt is an instruction: a series of tokens you send to an LLM. Every character, word, and punctuation mark is tokenized. Understanding this is the foundation: prompts are not magic strings, they are measured, budgeted, and optimized sequences of tokens.
1# Every prompt is a series of tokens2import tiktoken34enc = tiktoken.get_encoding("o200k_base")56system = "You are a helpful assistant."7user = "Fix the auth bug."89system_tokens = len(enc.encode(system))10user_tokens = len(enc.encode(user))1112print(f"System: {system_tokens} tokens")13print(f"User: {user_tokens} tokens")14print(f"Total: {system_tokens + user_tokens} tokens")
Context Window
The finite token budget
Every LLM has a finite context window, the maximum number of tokens it can process in a single request. Think of it as a desk: your prompt (instructions), conversation history, retrieved context, and response space must all fit. When the desk is full, something must go.
1# Context Window = Your Token Budget2import anthropic34client = anthropic.Anthropic()5WINDOW = 1_000_000 # Claude Sonnet 5.567count = client.messages.count_tokens(8 model="claude-sonnet-5-5",9 system="You are a helpful assistant.",10 messages=[{"role": "user", "content": "Fix the auth bug."}],11)12remaining = WINDOW - count.input_tokens13print(f"Budget remaining: {remaining:,} tokens")
Trimming (Last-N)
Delete the oldest, keep the recent
The simplest token optimization strategy. When the context window fills up, delete the oldest conversation turns and keep only the last N. Like tearing pages from the front of a notebook: fast and predictable, but you lose all early context.
1# Trimming: Last-N Strategy2def trim_history(messages, n=3):3 """Keep only the last N turns."""4 system = [m for m in messages if m['role'] == 'system']5 turns = [m for m in messages if m['role'] != 'system']6 return system + turns[-n * 2:]78# 20 messages → 6 (last 3 turns)9messages = trim_history(conversation, n=3)10print(f"Kept {len(messages)} messages")
Summarisation
Condense to save tokens
Instead of deleting old turns, summarise the entire conversation into a compact snapshot. You preserve the big picture but lose verbatim detail. The trade-off: an extra API call (more tokens spent) vs. richer context retention. This is token economy in action.
| Trimming (Last-N) | Summarisation | |
|---|---|---|
| Speed | Instant | Slow (LLM call) |
| Token cost | Free | Extra API call |
| Early context | Lost completely | Preserved (condensed) |
| Best for | Simple chatbots | Complex workflows |
| Risk | Amnesia | Detail loss |
1# Summarisation: Token Economy2import anthropic34def summarise_history(messages, client):5 """Condense conversation to save tokens."""6 history_text = "\n".join(7 f"{m['role']}: {m['content']}" for m in messages8 )9 response = client.messages.create(10 model="claude-sonnet-5-5",11 max_tokens=500,12 messages=[{13 "role": "user",14 "content": f"Summarise this conversation:\n{history_text}"15 }]16 )17 return response.content[0].text
Context Management
Token optimization and economy
Context management is the economy of tokens: deciding what goes into the window and what stays out. You allocate a token budget across system prompt, user message, retrieved context, and response space. Every token has a cost, and every token must earn its place.
1# Token Budget Manager2class TokenBudget:3 def __init__(self, limit=1_000_000): # Claude Sonnet 5.54 self.limit = limit5 self.allocations = {}67 def allocate(self, name, tokens):8 self.allocations[name] = tokens910 @property11 def remaining(self):12 used = sum(self.allocations.values())13 return self.limit - used1415 @property16 def utilization(self):17 return sum(self.allocations.values()) / self.limit
Context Engineering
IDE-driven, JIT context delivery
Modern IDEs don’t dump everything into the context window. They use just-in-time (JIT) context delivery, pulling in only the files, functions, and docs relevant to the current task. For long-horizon tasks spanning hundreds of tool calls, this intelligent context selection is essential.
| Dump Everything | JIT Context | |
|---|---|---|
| Strategy | Send all files | Pull relevant files on demand |
| Token usage | High (wasteful) | Low (efficient) |
| Quality | Diluted by noise | Focused signal |
| Best for | Small projects | Large codebases, long tasks |
| Example | Paste entire repo | IDE auto-includes imports |
Context Pollution
When tokens work against you
Not all tokens are equal. Irrelevant search results, stale tool outputs, and verbose error logs pollute the context, pushing out useful information and confusing the model. In a 200K token window processing 5 tickets, data from Ticket #1 clutters processing of Ticket #5.
1# Context Pollution Detection2def detect_pollution(messages):3 """Flag stale or redundant content."""4 stale = []5 for i, msg in enumerate(messages):6 if msg.get('tool_result'):7 age = len(messages) - i8 if age > 10: # older than 10 turns9 stale.append(i)10 print(f"Found {len(stale)} stale entries")11 return stale
Automatic Context Compaction
Server-side summarization
Server-side compaction summarizes earlier conversation history inside the API. Threshold mode (compact_20260112, beta compact-2026-01-12) runs when input tokens pass a trigger, 150K by default with a 50K minimum. On-demand mode (compaction: {type: "summarize"}, beta compact-2026-09-04) compacts when you ask. Anthropic’s cookbook measured 208K → 86K tokens (58.6%) on 5 support tickets with the older SDK helper compaction_control, which is now deprecated in the TypeScript and Ruby SDKs and removed in Python SDK v1.0.
1# Server-side compaction (threshold mode)2import anthropic34client = anthropic.Anthropic()56response = client.beta.messages.create(7 model="claude-sonnet-5-5",8 max_tokens=16000,9 betas=["compact-2026-01-12"],10 context_management={"edits": [{11 "type": "compact_20260112",12 "trigger": {"type": "input_tokens", "value": 150_000},13 }]},14 messages=messages,15)1617# Keep the compaction block: append the full content18messages.append({"role": "assistant", "content": response.content})
| Mode | When it runs | Beta header |
|---|---|---|
| compact_20260112 (threshold) | Input tokens pass the trigger (default 150K, minimum 50K) | compact-2026-01-12 |
| compaction: {type: "summarize"} (on-demand) | When your code asks for it | compact-2026-09-04 |
Sources: Compaction
Context Editing & Memory Tool
Auto-cleaner + filing cabinet
Context editing clears stale content by rule before the request reaches the model, with no secondary model involved: clear_tool_uses_20250919 removes old tool results and clear_thinking_20251015 removes old thinking blocks (beta context-management-2025-06-27). The memory tool provides persistent external storage (a filing cabinet) that survives across sessions. In Anthropic’s evaluation, context editing alone improved agentic task performance by 29%, editing plus memory by 39%, and editing cut token use by 84% in a 100-turn web search run.
| Context Editing | Memory Tool | |
|---|---|---|
| What it does | Removes stale clutter | Saves key facts permanently |
| Where | On the desk (context) | In the cabinet (external) |
| Token saving | 84% in a 100-turn web search run | Offloads to storage |
| Persistence | In-session only | Across all sessions |
| Performance | 29% better alone | 39% better combined with editing |
1# Context editing: clear old tool results by rule2response = client.beta.messages.create(3 model="claude-sonnet-5-5",4 max_tokens=16000,5 betas=["context-management-2025-06-27"],6 context_management={"edits": [{7 "type": "clear_tool_uses_20250919",8 "trigger": {"type": "input_tokens", "value": 100_000},9 "keep": {"type": "tool_uses", "value": 3},10 "clear_at_least": {"type": "input_tokens", "value": 10_000},11 "exclude_tools": ["memory"],12 }]},13 tools=tools,14 messages=messages,15)
Sources: Context editing · Context management results
Real-World: Customer Service
Compaction in production
A walkthrough based on Anthropic’s cookbook: 5 support tickets, 35+ tool calls, 208K tokens without compaction vs 86K with (measured with the deprecated compaction_control helper). The same ideas map to server-side compaction: a trigger for when to compact and instructions for what the summary must keep.
| Metric | No Compaction | With Compaction |
|---|---|---|
| Total turns | 37 | 26 |
| Input tokens | 204,416 | 82,171 |
| Output tokens | 4,422 | 4,275 |
| Total tokens | 208,838 | 86,446 |
| Compactions | N/A | 2 |
| Token savings | N/A | 122,392 (58.6%) |
1# Custom summary instructions for domain needs2context_management={"edits": [{3 "type": "compact_20260112",4 "trigger": {"type": "input_tokens", "value": 50_000},5 "instructions": (6 "Preserve: ticket IDs, categories, "7 "priorities, teams, outcomes. "8 "Discard: full KB articles, draft text."9 ),10}]}
Current API Levers
Caching, effort, thinking, tool search
Four settings decide most of what a Claude request costs today: prompt caching for repeated prefixes, effort for how hard the model works, adaptive thinking instead of fixed thinking budgets, and tool search so large tool catalogs stay out of context until needed.
| Prompt caching | Rate or limit |
|---|---|
| 5-minute cache write | 1.25x base input |
| 1-hour cache write | 2x base input |
| Cache read | 0.1x base input (0.05x on Opus 5.5 and Sonnet 5.5, 0.025x on Fable 5.1) |
| Breakpoints | Up to 4 per request |
| Minimum prefix | 512 tokens on most 5.x models; 1,024 on Opus 4.8, Sonnet 4.x, and Sonnet 5; 2,048 on Opus 4.7; 4,096 on Opus 4.5, Opus 4.6, and Haiku 4.5 |
| Invalidation | Changing tool definitions invalidates the whole cache |
| Lever | What to know |
|---|---|
| Effort | output_config.effort: low, medium, high, xhigh, max. Default high, except medium on Opus 5.5 and Haiku 5.5 |
| Adaptive thinking | thinking: {type: "adaptive"} replaces budget_tokens, which returns a 400 error on Opus 4.7 and later and on 5.x models |
| Tool search | Tools marked defer_loading: true stay out of context until the model searches for and loads them |
1# Effort, adaptive thinking, and caching in one request2response = client.messages.create(3 model="claude-opus-5-5",4 max_tokens=16000,5 thinking={"type": "adaptive"},6 output_config={"effort": "high"},7 cache_control={"type": "ephemeral"},8 system=STABLE_SYSTEM_PROMPT,9 messages=messages,10)
Sources: Prompt caching · Effort · Thinking · Tool search