@tokalator4.2.0…installs…downloads

Learn token economics

Short lessons on tokens, context windows and cost, with runnable code.

1

What Is a Prompt?

Instructions made of tokens

A prompt is an instruction: a series of tokens you send to an LLM. Every character, word, and punctuation mark is tokenized. Understanding this is the foundation: prompts are not magic strings, they are measured, budgeted, and optimized sequences of tokens.

Prompt = Instruction = Tokens
Tokenization
You
are
a
helpful
assistant
.
Fix
the
auth
bug
.
System prompt: 6 tokensUser message: 5 tokensTotal: 11 tokens
python
1# Every prompt is a series of tokens
2import tiktoken
3
4enc = tiktoken.get_encoding("o200k_base")
5
6system = "You are a helpful assistant."
7user = "Fix the auth bug."
8
9system_tokens = len(enc.encode(system))
10user_tokens = len(enc.encode(user))
11
12print(f"System: {system_tokens} tokens")
13print(f"User: {user_tokens} tokens")
14print(f"Total: {system_tokens + user_tokens} tokens")
2

Context Window

The finite token budget

Every LLM has a finite context window, the maximum number of tokens it can process in a single request. Think of it as a desk: your prompt (instructions), conversation history, retrieved context, and response space must all fit. When the desk is full, something must go.

Context Window: 1M tokens
System Prompt15%
Conversation25%
Retrieved Context35%
Response Space25%
Total: 1,000,000 tokensModel: Claude Sonnet 5.5Input cost: ~$1.50/req
python
1# Context Window = Your Token Budget
2import anthropic
3
4client = anthropic.Anthropic()
5WINDOW = 1_000_000 # Claude Sonnet 5.5
6
7count = client.messages.count_tokens(
8 model="claude-sonnet-5-5",
9 system="You are a helpful assistant.",
10 messages=[{"role": "user", "content": "Fix the auth bug."}],
11)
12remaining = WINDOW - count.input_tokens
13print(f"Budget remaining: {remaining:,} tokens")
3

Trimming (Last-N)

Delete the oldest, keep the recent

The simplest token optimization strategy. When the context window fills up, delete the oldest conversation turns and keep only the last N. Like tearing pages from the front of a notebook: fast and predictable, but you lose all early context.

Trimming: Last-N Strategy
Turn 1
Turn 2
Turn 3
Turn 4
Turn 5
Turn 6
Turn 7
Turn 8
Deleted (oldest 5)Kept (last 3)
python
1# Trimming: Last-N Strategy
2def trim_history(messages, n=3):
3 """Keep only the last N turns."""
4 system = [m for m in messages if m['role'] == 'system']
5 turns = [m for m in messages if m['role'] != 'system']
6 return system + turns[-n * 2:]
7
8# 20 messages → 6 (last 3 turns)
9messages = trim_history(conversation, n=3)
10print(f"Kept {len(messages)} messages")
4

Summarisation

Condense to save tokens

Instead of deleting old turns, summarise the entire conversation into a compact snapshot. You preserve the big picture but lose verbatim detail. The trade-off: an extra API call (more tokens spent) vs. richer context retention. This is token economy in action.

Summarisation: Snapshot Strategy
Full History
System prompt
User: setup project
AI: created files...
User: add auth
AI: implemented...
User: fix bug #42
AI: found issue...
~4,200 tokens
Summary
Project initialized with auth module. Bug #42 identified in token validation. Current focus: fixing edge case in refresh flow.
~180 tokens (96% reduction)
Trimming (Last-N)Summarisation
SpeedInstantSlow (LLM call)
Token costFreeExtra API call
Early contextLost completelyPreserved (condensed)
Best forSimple chatbotsComplex workflows
RiskAmnesiaDetail loss
python
1# Summarisation: Token Economy
2import anthropic
3
4def summarise_history(messages, client):
5 """Condense conversation to save tokens."""
6 history_text = "\n".join(
7 f"{m['role']}: {m['content']}" for m in messages
8 )
9 response = client.messages.create(
10 model="claude-sonnet-5-5",
11 max_tokens=500,
12 messages=[{
13 "role": "user",
14 "content": f"Summarise this conversation:\n{history_text}"
15 }]
16 )
17 return response.content[0].text
5

Context Management

Token optimization and economy

Context management is the economy of tokens: deciding what goes into the window and what stays out. You allocate a token budget across system prompt, user message, retrieved context, and response space. Every token has a cost, and every token must earn its place.

Token Budget Allocation
System Prompt12%
User Message8%
Retrieved Files45%
Conversation History20%
Response Reserve15%
85% allocated15% freeWarning: near limit
python
1# Token Budget Manager
2class TokenBudget:
3 def __init__(self, limit=1_000_000): # Claude Sonnet 5.5
4 self.limit = limit
5 self.allocations = {}
6
7 def allocate(self, name, tokens):
8 self.allocations[name] = tokens
9
10 @property
11 def remaining(self):
12 used = sum(self.allocations.values())
13 return self.limit - used
14
15 @property
16 def utilization(self):
17 return sum(self.allocations.values()) / self.limit
6

Context Engineering

IDE-driven, JIT context delivery

Modern IDEs don’t dump everything into the context window. They use just-in-time (JIT) context delivery, pulling in only the files, functions, and docs relevant to the current task. For long-horizon tasks spanning hundreds of tool calls, this intelligent context selection is essential.

JIT Context: Pull Only What You Need
✓ Current file
✓ Open tabs
✓ Import graph
✗ Git diff
✗ Test files
✗ Docs
3 sources active: IDE pulls context just-in-time, not all-at-once
Dump EverythingJIT Context
StrategySend all filesPull relevant files on demand
Token usageHigh (wasteful)Low (efficient)
QualityDiluted by noiseFocused signal
Best forSmall projectsLarge codebases, long tasks
ExamplePaste entire repoIDE auto-includes imports
7

Context Pollution

When tokens work against you

Not all tokens are equal. Irrelevant search results, stale tool outputs, and verbose error logs pollute the context, pushing out useful information and confusing the model. In a 200K token window processing 5 tickets, data from Ticket #1 clutters processing of Ticket #5.

Context Pollution: Token Waste
System prompt
500 tokens
User request
200 tokens
Stale tool output
8,200 tokens (wasted)
Old KB search
3,400 tokens (wasted)
Current task
600 tokens
Prev ticket draft
2,800 tokens (wasted)
Useful: 1,300 tokens (8%)Pollution: 14,400 tokens (92%)
python
1# Context Pollution Detection
2def detect_pollution(messages):
3 """Flag stale or redundant content."""
4 stale = []
5 for i, msg in enumerate(messages):
6 if msg.get('tool_result'):
7 age = len(messages) - i
8 if age > 10: # older than 10 turns
9 stale.append(i)
10 print(f"Found {len(stale)} stale entries")
11 return stale
8

Automatic Context Compaction

Server-side summarization

Server-side compaction summarizes earlier conversation history inside the API. Threshold mode (compact_20260112, beta compact-2026-01-12) runs when input tokens pass a trigger, 150K by default with a 50K minimum. On-demand mode (compaction: {type: "summarize"}, beta compact-2026-09-04) compacts when you ask. Anthropic’s cookbook measured 208K → 86K tokens (58.6%) on 5 support tickets with the older SDK helper compaction_control, which is now deprecated in the TypeScript and Ruby SDKs and removed in Python SDK v1.0.

Cookbook Results: 5 Tickets
Before
208K
37 turns
After
86K
58.6% saved
2
compaction events
58.6%
token reduction
26
turns (vs 37)
python
1# Server-side compaction (threshold mode)
2import anthropic
3
4client = anthropic.Anthropic()
5
6response = client.beta.messages.create(
7 model="claude-sonnet-5-5",
8 max_tokens=16000,
9 betas=["compact-2026-01-12"],
10 context_management={"edits": [{
11 "type": "compact_20260112",
12 "trigger": {"type": "input_tokens", "value": 150_000},
13 }]},
14 messages=messages,
15)
16
17# Keep the compaction block: append the full content
18messages.append({"role": "assistant", "content": response.content})
ModeWhen it runsBeta header
compact_20260112 (threshold)Input tokens pass the trigger (default 150K, minimum 50K)compact-2026-01-12
compaction: {type: "summarize"} (on-demand)When your code asks for itcompact-2026-09-04

Sources: Compaction

9

Context Editing & Memory Tool

Auto-cleaner + filing cabinet

Context editing clears stale content by rule before the request reaches the model, with no secondary model involved: clear_tool_uses_20250919 removes old tool results and clear_thinking_20251015 removes old thinking blocks (beta context-management-2025-06-27). The memory tool provides persistent external storage (a filing cabinet) that survives across sessions. In Anthropic’s evaluation, context editing alone improved agentic task performance by 29%, editing plus memory by 39%, and editing cut token use by 84% in a 100-turn web search run.

Context Editing vs Memory Tool
Context Editing
The Auto-Cleaner
84% token reduction
In-session only
Memory Tool
The Filing Cabinet
Prefs
Facts
Plans
Persists forever
Cross-session memory
Context EditingMemory Tool
What it doesRemoves stale clutterSaves key facts permanently
WhereOn the desk (context)In the cabinet (external)
Token saving84% in a 100-turn web search runOffloads to storage
PersistenceIn-session onlyAcross all sessions
Performance29% better alone39% better combined with editing
python
1# Context editing: clear old tool results by rule
2response = client.beta.messages.create(
3 model="claude-sonnet-5-5",
4 max_tokens=16000,
5 betas=["context-management-2025-06-27"],
6 context_management={"edits": [{
7 "type": "clear_tool_uses_20250919",
8 "trigger": {"type": "input_tokens", "value": 100_000},
9 "keep": {"type": "tool_uses", "value": 3},
10 "clear_at_least": {"type": "input_tokens", "value": 10_000},
11 "exclude_tools": ["memory"],
12 }]},
13 tools=tools,
14 messages=messages,
15)

Sources: Context editing · Context management results

10

Real-World: Customer Service

Compaction in production

A walkthrough based on Anthropic’s cookbook: 5 support tickets, 35+ tool calls, 208K tokens without compaction vs 86K with (measured with the deprecated compaction_control helper). The same ideas map to server-side compaction: a trigger for when to compact and instructions for what the summary must keep.

Customer Service Workflow: 5 Tickets
Per-ticket workflow (7 steps each)
Fetch
Classify
Research
Prioritize
Route
Draft
Complete
Linear token growth without compaction:
Turn 1204K tokens
MetricNo CompactionWith Compaction
Total turns3726
Input tokens204,41682,171
Output tokens4,4224,275
Total tokens208,83886,446
CompactionsN/A2
Token savingsN/A122,392 (58.6%)
python
1# Custom summary instructions for domain needs
2context_management={"edits": [{
3 "type": "compact_20260112",
4 "trigger": {"type": "input_tokens", "value": 50_000},
5 "instructions": (
6 "Preserve: ticket IDs, categories, "
7 "priorities, teams, outcomes. "
8 "Discard: full KB articles, draft text."
9 ),
10}]}
11

Current API Levers

Caching, effort, thinking, tool search

Four settings decide most of what a Claude request costs today: prompt caching for repeated prefixes, effort for how hard the model works, adaptive thinking instead of fixed thinking budgets, and tool search so large tool catalogs stay out of context until needed.

Prompt cachingRate or limit
5-minute cache write1.25x base input
1-hour cache write2x base input
Cache read0.1x base input (0.05x on Opus 5.5 and Sonnet 5.5, 0.025x on Fable 5.1)
BreakpointsUp to 4 per request
Minimum prefix512 tokens on most 5.x models; 1,024 on Opus 4.8, Sonnet 4.x, and Sonnet 5; 2,048 on Opus 4.7; 4,096 on Opus 4.5, Opus 4.6, and Haiku 4.5
InvalidationChanging tool definitions invalidates the whole cache
LeverWhat to know
Effortoutput_config.effort: low, medium, high, xhigh, max. Default high, except medium on Opus 5.5 and Haiku 5.5
Adaptive thinkingthinking: {type: "adaptive"} replaces budget_tokens, which returns a 400 error on Opus 4.7 and later and on 5.x models
Tool searchTools marked defer_loading: true stay out of context until the model searches for and loads them
python
1# Effort, adaptive thinking, and caching in one request
2response = client.messages.create(
3 model="claude-opus-5-5",
4 max_tokens=16000,
5 thinking={"type": "adaptive"},
6 output_config={"effort": "high"},
7 cache_control={"type": "ephemeral"},
8 system=STABLE_SYSTEM_PROMPT,
9 messages=messages,
10)

Sources: Prompt caching · Effort · Thinking · Tool search