Skip to content
thu-nicsPublic

About

Official implementation of paper "Cache Engineering for Agents"

Resources

Stars

9 stars

Watchers

0 watching

Forks

Latest commit

 

History

13 Commits

Folders and files

Repository files navigation

AKV animated memory bot

AKV

Cache Engineering for Agents

🌐 Project Page

AKV brings KV cache engineering to LLM agents. It combines agent-side policies for deciding what to drop and when to compact positions with an SGLang-based inference system that executes those operations while reusing compatible KV cache.

  • akv.agent is the algorithm component that decides which messages to drop from the KV cache and when to reposition retained tokens.
  • akv.system executes cache operations through SGLang's OpenAI-compatible Chat Completions API.

Environment Setup

  1. Install Rust and Cargo if they are not already available:
curl -o install-rust.sh https://sh.rustup.rs
sh install-rust.sh -y
source "$HOME/.cargo/env"
  1. Create the akv environment and install the agent and system together:
git clone https://github.com/thu-nics/AKV.git
cd AKV
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e .

The environment is stored in .venv to avoid conflicting with the akv/ source directory. See the system README and the agent README for dependencies.

Quick Start

  1. Launch the system
python -m akv.system.launch_server \
  --model-path Qwen/Qwen3-0.6B \
  --port 8000

For additional launch options, see the system README.

  1. Send a simple request
import requests

response = requests.post(
    "http://localhost:8000/v1/chat/completions",
    json={
        "model": "Qwen/Qwen3-0.6B",
        "messages": [
            {"role": "system", "content": "You are a helpful assistant."},
            {"role": "user", "content": "The capital of France is Paris."},
            {"role": "assistant", "content": "I have read the note."},
            {"role": "user", "content": "What is the capital of France?"},
        ],
        "drop_message": {"3": [1, 2]},
        "reposition": [3],
        "temperature": 0,
        "max_tokens": 2048,
    },
    timeout=120,
)
response.raise_for_status()
print(response.json())

Message indices are zero-based. After processing message 3, the server drops the KV of messages 1 and 2 and compacts the retained positions before generating the response.

  1. Apply an AKV policy

Apply the policy to a Chat Completions payload containing the full history, tools, and generation options, without existing cache instructions:

from akv.agent import CachePolicy, CacheState

policy = CachePolicy(
    rolling_drop_target="tool",
    drop_after_call=1,
    drop_interval_calls=1,
    drop_target_responses=12,
    reposition_after_call=4,
    reposition_interval_calls=4,
)
state = CacheState()

prepared, next_state = policy.process(payload, state=state, api_style="sglang")

Send prepared as the request body. After a successful response, reuse next_state with the full updated history for the next request. See the agent README and complete agent demo for usage examples.

Usage

System: Drop and Reposition

Request field Usage
drop_message Maps a trigger message index to the message indices whose KV should be dropped after that trigger.
drop_rule Selects a structured Drop rule. message_drop is equivalent to drop_message; thinking_drop removes historical assistant reasoning KV while retaining the final answer. Do not combine drop_rule with drop_message.
reposition Lists message indices after which to compact retained KV positions. Optional and independent of Drop; indices must be unique and strictly increasing.

Drop alone preserves the surviving positions; Reposition removes the gaps. Thinking Drop acts on historical reasoning and does not disable thinking generation for the new reply.

Context requests report cached_tokens, drop_skipped_tokens, and repos_tokens in usage.prompt_tokens_details, including zero values. See the system README for rule details and usage accounting.

Agent: policy configuration

CachePolicy supports rules or agent-selected actions. CachePolicy() applies no automatic operations.

ConfigurationParameterUsage
What to drop rolling_drop_target Select tool, assistant, or both for deletion based on the retention window.
When to drop tools drop_after_call=N First check after complete tool batch N.
drop_interval_calls=M Repeat every M batches; requires drop_after_call. Omit for a single check.
drop_trigger_tokens=T Check when the visible prompt exceeds T tokens.
How much to retain drop_target_responses=K Window size in individual tool results.
drop_target_tokens=U Drop oldest tool results toward a visible prompt length of U tokens.
When to reposition reposition_after_call=N First reposition after complete tool batch N.
reposition_interval_calls=M Repeat every M batches; requires reposition_after_call. Omit for a single operation.
reposition_trigger_tokens=T Reposition when token position length exceeds T and dropped content leaves gaps.
Token counting token_counter Required for token thresholds and token retention targets.

For each operation, choose either batch or token triggers. Choose one retention target; a token target must be below the token Drop trigger when both are set. The rolling example above checks after every batch and drops only when more than 12 tool results are retained.

Pass options to process to override rules for one request, or actions to supply agent-selected message IDs and a reposition decision. See the agent README for examples and configuration details.

Supported Models

Drop/Reposition is currently enabled for these text-model architectures:

Model family Architecture
Qwen QWenLMHeadModel
Qwen1.5, Qwen2, Qwen2.5 Qwen2ForCausalLM, Qwen2MoeForCausalLM
Qwen3, including MoE variants and AgenticQwen Qwen3ForCausalLM, Qwen3MoeForCausalLM
GPT-OSS GptOssForCausalLM
MiniMax-M2 family, including M2.7 MiniMaxM2ForCausalLM

See the system README for runtime requirements and compatibility details.

About

Official implementation of paper "Cache Engineering for Agents"

Resources

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages