AKV brings KV cache engineering to LLM agents. It combines agent-side policies for deciding what to drop and when to compact positions with an SGLang-based inference system that executes those operations while reusing compatible KV cache.
akv.agentis the algorithm component that decides which messages to drop from the KV cache and when to reposition retained tokens.akv.systemexecutes cache operations through SGLang's OpenAI-compatible Chat Completions API.
- Install Rust and Cargo if they are not already available:
curl -o install-rust.sh https://sh.rustup.rs
sh install-rust.sh -y
source "$HOME/.cargo/env"- Create the
akvenvironment and install the agent and system together:
git clone https://github.com/thu-nics/AKV.git
cd AKV
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e .The environment is stored in .venv to avoid conflicting with the akv/ source directory. See the system README and the agent README for dependencies.
- Launch the system
python -m akv.system.launch_server \
--model-path Qwen/Qwen3-0.6B \
--port 8000For additional launch options, see the system README.
- Send a simple request
import requests
response = requests.post(
"http://localhost:8000/v1/chat/completions",
json={
"model": "Qwen/Qwen3-0.6B",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "The capital of France is Paris."},
{"role": "assistant", "content": "I have read the note."},
{"role": "user", "content": "What is the capital of France?"},
],
"drop_message": {"3": [1, 2]},
"reposition": [3],
"temperature": 0,
"max_tokens": 2048,
},
timeout=120,
)
response.raise_for_status()
print(response.json())Message indices are zero-based. After processing message 3, the server drops the KV of messages 1 and 2 and compacts the retained positions before generating the response.
- Apply an AKV policy
Apply the policy to a Chat Completions payload containing the full history, tools, and generation options, without existing cache instructions:
from akv.agent import CachePolicy, CacheState
policy = CachePolicy(
rolling_drop_target="tool",
drop_after_call=1,
drop_interval_calls=1,
drop_target_responses=12,
reposition_after_call=4,
reposition_interval_calls=4,
)
state = CacheState()
prepared, next_state = policy.process(payload, state=state, api_style="sglang")Send prepared as the request body. After a successful response, reuse next_state with the full updated history for the next request. See the agent README and complete agent demo for usage examples.
| Request field | Usage |
|---|---|
drop_message |
Maps a trigger message index to the message indices whose KV should be dropped after that trigger. |
drop_rule |
Selects a structured Drop rule. message_drop is equivalent to drop_message; thinking_drop removes historical assistant reasoning KV while retaining the final answer. Do not combine drop_rule with drop_message. |
reposition |
Lists message indices after which to compact retained KV positions. Optional and independent of Drop; indices must be unique and strictly increasing. |
Drop alone preserves the surviving positions; Reposition removes the gaps. Thinking Drop acts on historical reasoning and does not disable thinking generation for the new reply.
Context requests report cached_tokens, drop_skipped_tokens, and repos_tokens in usage.prompt_tokens_details, including zero values. See the system README for rule details and usage accounting.
CachePolicy supports rules or agent-selected actions. CachePolicy() applies no automatic operations.
| Configuration | Parameter | Usage |
|---|---|---|
| What to drop | rolling_drop_target |
Select tool, assistant, or both for deletion based on the retention window. |
| When to drop tools | drop_after_call=N |
First check after complete tool batch N. |
drop_interval_calls=M |
Repeat every M batches; requires drop_after_call. Omit for a single check. |
|
drop_trigger_tokens=T |
Check when the visible prompt exceeds T tokens. | |
| How much to retain | drop_target_responses=K |
Window size in individual tool results. |
drop_target_tokens=U |
Drop oldest tool results toward a visible prompt length of U tokens. | |
| When to reposition | reposition_after_call=N |
First reposition after complete tool batch N. |
reposition_interval_calls=M |
Repeat every M batches; requires reposition_after_call. Omit for a single operation. |
|
reposition_trigger_tokens=T |
Reposition when token position length exceeds T and dropped content leaves gaps. | |
| Token counting | token_counter |
Required for token thresholds and token retention targets. |
For each operation, choose either batch or token triggers. Choose one retention target; a token target must be below the token Drop trigger when both are set. The rolling example above checks after every batch and drops only when more than 12 tool results are retained.
Pass options to process to override rules for one request, or actions to supply agent-selected message IDs and a reposition decision. See the agent README for examples and configuration details.
Drop/Reposition is currently enabled for these text-model architectures:
| Model family | Architecture |
|---|---|
| Qwen | QWenLMHeadModel |
| Qwen1.5, Qwen2, Qwen2.5 | Qwen2ForCausalLM, Qwen2MoeForCausalLM |
| Qwen3, including MoE variants and AgenticQwen | Qwen3ForCausalLM, Qwen3MoeForCausalLM |
| GPT-OSS | GptOssForCausalLM |
| MiniMax-M2 family, including M2.7 | MiniMaxM2ForCausalLM |
See the system README for runtime requirements and compatibility details.
