@tokalator4.2.0…installs…downloads

Papers

Research on context engineering, prompt caching and agentic coding, searchable.

arXivContext Management

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

A widespread practice in software development is to tailor coding agents to repositories using context files, such as AGENTS.md. Although this practice is strongly encouraged by agent developers,...

Thibaud Gloaguen, Niels Mündler-Sasahara +3
cs.SEcs.AI
arXivContext Management

Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models

Large language model (LLM) applications such as agents and domain-specific reasoning increasingly rely on context adaptation: modifying inputs with instructions, strategies, or evidence, rather than...

Qizheng Zhang, Changran Hu +11
cs.LGcs.AIcs.CL
arXivCaching

Don't Break the Cache: An Evaluation of Prompt Caching for Long-Horizon Agentic Tasks

Recent advancements in Large Language Model (LLM) agents have enabled complex multi-turn agentic tasks requiring extensive tool calling, where conversations can span dozens of API calls with...

Elias Lumer, Faheem Nizar +5
cs.CL
arXivContext Management

A Survey of Context Engineering for Large Language Models

The performance of Large Language Models (LLMs) is fundamentally determined by the contextual information provided during inference. This survey introduces Context Engineering, a formal discipline...

Lingrui Mei, Jiayu Yao +13
cs.CL
arXivContext Management

The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management

Large Language Model (LLM)-based agents solve complex tasks through iterative reasoning, exploration, and tool-use, a process that can result in long, expensive context histories. While...

Tobias Lindenbauer, Igor Slinko +3
cs.SEcs.AI
arXivContext Management

Context Length Alone Hurts LLM Performance Despite Perfect Retrieval

Large language models (LLMs) often fail to scale their performance on long-context tasks performance in line with the context lengths they support. This gap is commonly attributed to retrieval...

Yufeng Du, Minyang Tian +8
cs.CLcs.AI
arXivContext Management

Agent READMEs: An Empirical Study of Context Files for Agentic Coding

Agentic coding tools receive goals written in natural language, break them down into specific tasks, and write or execute code with minimal human intervention. Central to this process are agent...

Worawalan Chatlatanagulchai, Hao Li +9
cs.SE
arXivContext Management

Context Engineering 2.0: The Context of Context Engineering

Karl Marx once wrote that ``the human essence is the ensemble of social relations'', suggesting that individuals are not isolated entities but are fundamentally shaped by their interactions with...

Qishuo Hua, Lyumanshan Ye +7
cs.AIcs.CL
arXivContext Management

Monadic Context Engineering

The proliferation of Large Language Models (LLMs) has catalyzed a shift towards autonomous agents capable of complex reasoning and tool use. However, current agent architectures are frequently...

Yifan Zhang, Yang Yuan +2
cs.AIcs.CLcs.FL
arXivContext Management

The Root Theorem of Context Engineering

Every system that maintains a large language model conversation beyond a single session faces two inescapable constraints: the context window is finite, and information quality degrades with...

Borja Odriozola Schick
cs.CCcs.CLcs.HC+1
arXivContext Management

CEDAR: Context Engineering for Agentic Data Science

We demonstrate CEDAR, an application for automating data science (DS) tasks with an agentic setup. Solving DS problems with LLMs is an underexplored area that has immense market value. The challenges...

Rishiraj Saha Roy, Chris Hinze +2
cs.LGcs.AI
arXivContext Management

Meta Context Engineering via Agentic Skill Evolution

The operational efficacy of large language models relies heavily on their inference-time context. This has established Context Engineering (CE) as a formal discipline for optimizing these inputs....

Haoran Ye, Xuning He +3
cs.AIcs.NE
arXivContext Management

TRACE: TRajectory Attribution for Automated Context Engineering

Production AI agents fail when their context sources -- system prompts, knowledge bases, tool descriptions, and procedural skills -- contain errors or gaps. Current maintenance relies on manual log...

Yikai Zhao, Pradeep Kumar Misra +1
cs.AIcs.LG
arXivContext Management

Ref-Long: Benchmarking the Long-context Referencing Capability of Long-context Language Models

Long-context language models (LCLMs) have exhibited impressive capabilities in long-context understanding tasks. Among these, long-context referencing -- a crucial task that requires LCLMs to...

Junjie Wu, Gefei Gu +3
cs.CL
arXivContext Management

Lost in the Middle: How Language Models Use Long Contexts

While recent language models have the ability to take long contexts as input, relatively little is known about how well they use longer context. We analyze the performance of language models on two...

Nelson F. Liu, Kevin Lin +5
cs.CL
arXivContext Management

LongGenBench: Long-context Generation Benchmark

Current long-context benchmarks primarily focus on retrieval-based tests, requiring Large Language Models (LLMs) to locate specific information within extensive input contexts, such as the...

Xiang Liu, Peijie Dong +2
cs.CLcs.AI
arXivContext Management

Towards Long Context Hallucination Detection

Large Language Models (LLMs) have demonstrated remarkable performance across various tasks. However, they are prone to contextual hallucination, generating information that is either unsubstantiated...

Siyi Liu, Kishaloy Halder +7
cs.CLcs.AI
arXivContext Management

DocFinQA: A Long-Context Financial Reasoning Dataset

For large language models (LLMs) to be effective in the financial domain -- where each decision can have a significant impact -- it is necessary to investigate realistic tasks and data. Financial...

Varshini Reddy, Rik Koncel-Kedziorski +4
cs.CLcs.AI
arXivContext Management

Evaluating Zero-Shot Long-Context LLM Compression

This study evaluates the effectiveness of zero-shot compression techniques on large language models (LLMs) under long-context. We identify the tendency for computational errors to increase under...

Chenyu Wang, Yihan Wang +1
cs.CLcs.AI
arXivCaching

CacheProbe: Auditing Prompt Cache Isolation in Gateway APIs

Over the past year, prompt caching in Large Language Models (LLMs) has become increasingly more popular across inference APIs. Prompt caching helps save precious compute resources and speeds up...

Ryan Fahey
cs.CRcs.LG
arXivCaching

Auditing Prompt Caching in Language Model APIs

Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than non-cached prompts. These timing differences introduce the risk of...

Chenchen Gu, Xiang Lisa Li +3
cs.CLcs.CRcs.LG
arXivCaching

Prompt Cache: Modular Attention Reuse for Low-Latency Inference

We present Prompt Cache, an approach for accelerating inference for large language models (LLM) by reusing attention states across different LLM prompts. Many input prompts have overlapping text...

In Gim, Guojun Chen +4
cs.CLcs.AI
arXivCaching

Efficient Prompt Caching via Embedding Similarity

Large language models (LLMs) have achieved huge success in numerous natural language process (NLP) tasks. However, it faces the challenge of significant resource consumption during inference. In this...

Hanlin Zhu, Banghua Zhu +1
cs.CLcs.LG
arXivCaching

Tail-Optimized Caching for LLM Inference

Prompt caching is critical for reducing latency and cost in LLM inference: OpenAI and Anthropic report up to 50-90% cost savings through prompt reuse. Despite its widespread success, little is known...

Wenxin Zhang, Yueying Li +2
eess.SY
arXivToken Optimization

The Token Efficiency Index: A Peer-Benchmarked Composite Indicator for AI Token Efficiency

As artificial intelligence (AI) adoption accelerates across tech giants, AI-native startups, and non-technical organizations alike, a deceptively simple question remains hard to answer: is that...

Caden Wong, Vikram Das +1
cs.PFcs.CE
arXivToken Optimization

A*-Decoding: Token-Efficient Inference Scaling

Inference-time scaling has emerged as a powerful alternative to parameter scaling for improving language model performance on complex reasoning tasks. While existing methods have shown strong...

Giannis Chatziveroglou
cs.AI
arXivToken Optimization

Token-Efficient RL for LLM Reasoning

We propose reinforcement learning (RL) strategies tailored for reasoning in large language models (LLMs) under strict memory and compute limits, with a particular focus on compatibility with LoRA...

Alan Lee, Harry Tong
cs.LGcs.AI
arXivToken Optimization

SkillReducer: Optimizing LLM Agent Skills for Token Efficiency

LLM-based coding agents rely on \emph{skills}, pre-packaged instruction sets that extend agent capabilities, yet every token of skill content injected into the context window incurs both monetary...

Yudong Gao, Zongjie Li +4
cs.SE
arXivToken Optimization

Token-Efficient Change Detection in LLM APIs

Remote change detection in LLMs is a difficult problem. Existing methods are either too expensive for deployment at scale, or require initial white-box access to model weights or grey-box access to...

Timothée Chauvin, Clément Lalanne +4
cs.LGcs.CR
arXivToken Optimization

Token-Efficient Leverage Learning in Large Language Models

Large Language Models (LLMs) have excelled in various tasks but perform better in high-resource scenarios, which presents challenges in low-resource scenarios. Data scarcity and the inherent...

Yuanhao Zeng, Min Wang +2
cs.CLcs.AIcs.LG
arXivAgents

DeepCode: Open Agentic Coding

Recent advances in large language models (LLMs) have given rise to powerful coding agents, making it possible for code assistants to evolve into code engineers. However, existing methods still face...

Zongwei Li, Zhonghang Li +3
cs.SEcs.AI
arXivAgents

Self-Evolving Coding Agents

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches....

Hao Zhou, Haichuan Hu +6
cs.SE
arXivAgents

OpenGame: Open Agentic Coding for Games

Game development sits at the intersection of creative design and intricate software engineering, demanding the joint orchestration of game engines, real-time loops, and tightly coupled state across...

Yilei Jiang, Jinyuan Hu +9
cs.SE
arXivAgents

Humanize: Judgement Engineering for Agentic Coding

Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent...

Sihao Liu, Ligeng Zhu +7
cs.AIcs.CY
arXivAgents

Bayesian control for coding agents

Modern coding agents pair LLM generators with various tools, including cheap diagnostics and expensive verifiers. The tool-use decisions are typically governed by orchestrators that often use fixed...

Theodore Papamarkou, Vladislav Smirnov +5
cs.AIcs.CL
arXivAgents

Scaling Test-Time Compute for Agentic Coding

Test-time scaling has become a powerful way to improve large language models. However, existing methods are best suited to short, bounded outputs that can be directly compared, ranked or refined....

Joongwon Kim, Wannan Yang +14
cs.SEcs.AIcs.CL+1
arXivContext Management

Context Rot in AI-Assisted Software Development: Repurposing Documentation Consistency for AI Configuration Artifacts

Developers increasingly provide AI coding assistants with persistent context through configuration files such as CLAUDE.md, AGENTS.md, and .cursorrules. These files describe code elements,...

Christoph Treude, Sebastian Baltes
cs.SEcs.AI
arXivContext Management

Diagnosing and Mitigating Context Rot in Long-horizon Search

Extensive context has become the norm as Large Language Models (LLMs) are increasingly deployed in long-horizon search tasks. The concern that increasing context length degrades model capabilities,...

Shijie Xia, Yikun Wang +2
cs.IRcs.AIcs.CL
arXivContext Management

When and How Context Rot Appears in Coding Agents: A White-Box Study of Agent Skills in Code Auditing

Agent Skills package procedural instructions and checks for use by general-purpose agents, but loading a skill does not guarantee that every requirement remains active throughout a long tool-using...

Yue Xue
cs.SE
arXivContext Management

Classifier Context Rot: Monitor Performance Degrades with Context Length

Monitoring coding agents for dangerous behavior using language models requires classifying transcripts that often exceed 500K tokens, but prior agent monitoring benchmarks rarely contain transcripts...

Sam Martin, Fabien Roger
cs.AI
arXivContext Management

The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus

LLMs are increasingly used as general-purpose reasoners, but long inputs remain bottlenecked by a fixed context window. Recursive Language Models (RLMs) address this by externalising the prompt and...

Amartya Roy, Rasul Tutunov +3
cs.LGcs.AI
arXivContext Management

DISCO: Distributed Long Context Scaling with Grounding-Reasoning Disaggregation

While Large Language Models (LLMs) advertise million-token context windows, reasoning quality often collapses as inputs grow -- a phenomenon termed context rot. This failure stems from a structural...

Guanzheng Chen, Viet Dac Lai +6
cs.CL
arXivContext Management

Kinetic theory for Transformers and the lost-in-the-middle phenomenon

We study causal self-attention dynamics -- a toy model for decoder Transformers -- which we interpret as a non-exchangeable interacting particle system. Adapting cumulant expansions to the triangular...

Mitia Duerinckx, Borjan Geshkovski +1
math.APcs.LGmath.PR
arXivContext Management

Lost in the Middle at Birth: An Exact Theory of Transformer Position Bias

The ``Lost in the Middle'' phenomenon -- a U-shaped performance curve where LLMs retrieve well from the beginning and end of a context but fail in the middle -- is widely attributed to learned...

Borun D Chowdhury
cs.LGcs.AIcs.CL
arXivContext Management

What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations

The diminishing ability of large language models (LLMs) to effectively utilize long-range context-the "lost-in-the-middle" phenomenon-poses a significant challenge in retrieval-based LLM...

Mihir Gupte, Eshan Dixit +2
cs.CLcs.AI
arXivContext Management

Lost in the Middle: An Emergent Property from Information Retrieval Demands in LLMs

The performance of Large Language Models (LLMs) often degrades when crucial information is in the middle of a long context, a "lost-in-the-middle" phenomenon that mirrors the primacy and recency...

Nikolaus Salvatore, Hao Wang +1
cs.LGq-bio.NC
arXivContext Management

An Adjoint-Sensitivity Framework for Lost-in-the-Middle Phenomena in Causal Residual Transformers

We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions. The...

Cheng Huan, Hongwei Yuan
stat.MLcs.LGmath.OC
arXivToken Optimization

Prompt Compression via Activation Aggregation

Large language models process prompts by propagating activations through dozens of layers before generating a response. We ask whether the task-relevant information contained in an instruction prompt...

Thibaud Ardoin, Semira Einsele +2
cs.CLcs.LG
arXivToken Optimization

Discrete Prompt Compression with Reinforcement Learning

Compressed prompts aid instruction-tuned language models (LMs) in overcoming context window limitations and reducing computational costs. Existing methods, which primarily based on training...

Hoyoun Jung, Kyung-Joong Kim
cs.CLcs.AI
arXivToken Optimization

Parse Trees Guided LLM Prompt Compression

Offering rich contexts to Large Language Models (LLMs) has shown to boost the performance in various tasks, but the resulting longer prompt would increase the computational cost and might exceed the...

Wenhao Mao, Chengbin Hou +4
cs.CLcs.AI
arXivToken Optimization

ProCut: LLM Prompt Compression via Attribution Estimation

In large-scale industrial LLM systems, prompt templates often expand to thousands of tokens as teams iteratively incorporate sections such as task instructions, few-shot examples, and heuristic rules...

Zhentao Xu, Fengyi Li +2
cs.CLcs.LG
arXivToken Optimization

EFPC: Towards Efficient and Flexible Prompt Compression

The emergence of large language models (LLMs) like GPT-4 has revolutionized natural language processing (NLP), enabling diverse, complex tasks. However, extensive token counts lead to high...

Yun-Hao Cao, Yangsong Wang +5
cs.CLcs.AI
arXivToken Optimization

Better Prompt Compression Without Multi-Layer Perceptrons

Prompt compression is a promising approach to speeding up language model inference without altering the generative model. Prior works compress prompts into smaller sequences of learned tokens using...

Edouardo Honig, Andrew Lizarraga +2
cs.CLcs.LG
Claude CookbookCaching

Prompt caching with the Claude API

Prompt caching lets you store and reuse context within your prompts, reducing latency by >2x and costs by up to 90% for repetitive tasks.

tokenspromptscaching
Claude CookbookCaching

Speculative Prompt Caching

This cookbook demonstrates "Speculative Prompt Caching" - a pattern that reduces time-to-first-token (TTFT) by warming up the cache while users are still formulating their queries.

tokenspromptscaching+1
Claude CookbookContext Management

Session Memory Compaction

Long-running conversations with Claude can exceed context limits, causing loss of important information. Whether you're building a coding assistant, creative writing tool, or customer service agent,...

tokenspromptscaching
Claude CookbookContext Management

Automatic Context Compaction

Long-running agentic tasks can often exceed context limits. Tool heavy workflows or long conversations quickly consume the token context window. In [Effective Context Engineering for AI...

tokenspromptscaching
Claude CookbookContext Management

Context Engineering for AI Agents: Memory vs. Compaction vs. Tool Clearing

A common challenge when building long-horizon agents is managing context. Tool results, the model's own reasoning, and user messages all accumulate, and eventually you either hit the token limit or...

tokenspromptscaching
Claude CookbookContext Management

Context Editing & Memory for Long-Running Agents

AI agents that run across multiple sessions or handle long-running tasks face two key challenges: they lose learned patterns between conversations, and context windows fill up during extended...

tokenspromptscaching
Claude CookbookToken Optimization

Cost Optimization on the Claude API

AI adoption is starting to take a familiar shape: you build on AI to stay at the intelligence frontier, but scaling your product makes token costs unsustainable. The zeitgeist refers to this...

tokenspromptscaching+1
Claude CookbookContext Management

Extended Thinking

This notebook demonstrates how to use Claude 3.7 Sonnet's extended thinking feature with various examples and edge cases.

tokenspromptsrate-limits+1
OpenAI CookbookGeneral

How to work with large language models

The magic of large language models is that by being trained to minimize this prediction error over vast quantities of text, the models end up learning concepts useful for these predictions. For...

tokensprompts
OpenAI CookbookPrompt Engineering

Techniques to improve reliability

When GPT-3 fails on a task, what should you do?

tokensprompts
OpenAI CookbookPrompt Engineering

Related resources from around the web

People are writing great tools and papers for improving outputs from GPT. Here are some cool ones we've seen:

prompts
OpenAI CookbookToken Optimization

How to count tokens with tiktoken

Given a text string (e.g., `"tiktoken is great!"`) and an encoding (e.g., `"cl100k_base"`), a tokenizer can split the text string into a list of tokens (e.g., `["t", "ik", "token", " is", " great",...

tokenspromptsembeddings
OpenAI CookbookGeneral

How to stream completions

By default, when you request a completion from the OpenAI, the entire completion is generated before being sent back in a single response.

tokensstreaming
Gemini CookbookCaching

Gemini API: Context Caching Quickstart

This notebook introduces context caching with the Gemini API and provides examples of interacting with the Apollo 11 transcript using the Python SDK. For a more comprehensive look, check out [the...

tokenspromptscaching
Gemini CookbookToken Optimization

Gemini API: All about tokens

An understanding of tokens is central to using the Gemini API. This guide will provide a interactive introduction to what tokens are and how they are used in the Gemini API.

tokenspromptscaching
Gemini CookbookGeneral

Use Gemini thinking

All Gemini models are trained to do a [thinking process](https://ai.google.dev/gemini-api/docs/thinking-mode) (or reasoning) before getting to a final answer. As a result, those models usually get...

tokenspromptsstreaming
Gemini CookbookContext Management

Gemini API: Adding context information

While LLMs are trained extensively on various documents and data, the LLM does not know everything. New information or information that is not easily accessible cannot be known by the LLM, unless it...

prompts
Gemini CookbookGeneral

Cost Estimation and Health Monitoring with Gemini

Cost observability is key for scaling Gemini API applications. Without tracking token usage and costs, you can't optimize your spending or identify expensive operations.

tokensprompts
Built-inContext Management

Progressive Disclosure

Instead of loading an entire codebase (which would immediately overwhelm the attention budget), modern agents use JIT context. The assistant dynamically loads only the necessary data at runtime.

contextjitoptimization
Built-inContext Management

Lightweight Identifiers

The assistant maintains references (file paths, stored queries) and dynamically loads only the necessary data at runtime using tools like grep, head, or tail.

contextreferencesefficiency
Built-inContext Management

Compaction

When a session nears its token limit, the assistant summarizes critical details, such as architectural decisions and unresolved bugs, while discarding redundant tool outputs.

contextcompressionlong-horizon
Built-inContext Management

Tool Result Clearing

A light touch form of compaction where the raw results of previous tool calls (like long terminal outputs) are cleared to save space.

contexttoolsoptimization
Built-inContext Management

Structured Note-taking

The agent may maintain an external NOTES.md or a to-do list to track dependencies and progress across thousands of steps, which it can read back into its context after a reset.

contextpersistencenotes
Built-inContext Management

Distractors

Files or code snippets that are topically related to the query but do not contain the answer can cause the model to lose focus or hallucinate.

contextpollutionrelevance
Built-inContext Management

Context Rot

As more tokens are added, the model's ability to accurately retrieve needles of information from the haystack of the codebase decreases.

contextdegradationtokens
Built-inPrompt Engineering

XML Tagging

Use tags like <background_information>, <tool_guidance>, <constraints> to clearly separate different types of instructions in system prompts.

promptsxmlstructure
Built-inToken Optimization

High-Signal Tokens

The objective is to provide the smallest possible set of high-signal tokens that maximize the likelihood of the correct code generation.

tokensoptimizationquality
Built-inContext Management

Structural Patterns

Research suggests that models often perform better on shuffled or unstructured context than on logically structured haystacks, impacting how they process long files.

contextstructureresearch
Built-inArchitecture

Agent Skills

Reusable packages of domain expertise defined in SKILL.md files that provide specialized AI agent capabilities. Introduced as GA in VS Code 1.109, skills can be invoked as slash commands or loaded...

skillsagentsvscode+1
Built-inArchitecture

Agent Hooks

Deterministic shell commands that execute at key lifecycle points during agent sessions. Unlike instructions, hooks run code with guaranteed outcomes for security policies, quality checks, or audit...

hooksagentslifecycle+1
Built-inArchitecture

Agent Orchestration

A multi-agent pattern where specialized subagents collaborate on complex tasks, each operating in its own dedicated context window. Provides context efficiency, specialization with different models,...

orchestrationmulti-agentsubagent+1
Built-inContext Management

Message Steering

An agent interaction pattern where follow-up messages redirect a running agent request. The agent yields after the active tool execution and processes the new message. Alternatives include request...

agentssteeringqueueing+1
Built-inArchitecture

Terminal Sandboxing

A security mechanism restricting file system and network access for agent-executed terminal commands. Sandboxed commands have read/write access only to the workspace directory, and network access can...

securitysandboxterminal+1
Built-inToken Economics

Thinking Tokens

Tokens generated during a model's internal reasoning process before producing a visible response. Thinking tokens consume context budget but improve quality on complex tasks. On current Claude...

thinkingreasoningtokens+1
Built-inArchitecture

MCP Server (Model Context Protocol)

A local stdio process that exposes tools to Claude Code and other MCP-capable agents. Tokalator's MCP server (tokalator-mcp) provides four tools: count_tokens, estimate_budget, preview_turn, and...

mcpclaude-codetools+2
Built-inToken Optimization

CLI Token Counter

A standalone command-line tool for counting tokens and checking context budgets outside of VS Code. Tokalator ships a CLI binary (tokalator count, budget, preview, models) for SSH sessions, CI...

cliterminaltokens+2

Last fetched: · 90 articles