Radix tree attention cache
SGLang Radix Tree attention memory reuse for complex multi-turn prompt contexts.
problem it solves
Prevents redundant prefill attention calculation across branching multi-turn chat sessions and long-context prompts.
What it does
What it does: Reuses Radix Tree attention nodes across branching multi-turn chat prompts.
What it watches: Prefix tree matches on self-hosted SGLang serving clusters.
When it triggers: When incoming prompt tokens share prefix nodes with active Radix Tree branches.
The Action: Skips prefill computation for matching prefix branches, drastically reducing TTFT.
How it recovers: Non-matching branch extensions undergo standard prefill processing.
What we need from you
Setup required before this can be enabled
Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.SGLang, or vLLM built with RadixTree memory management and radix attention left enabled.
- Self-hosted SGLang or vLLM deploymentrequired
Requires a model server with a RadixTree memory manager.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Your model server re-computes full prefill attention on every request. | No radix_cache_attention stage recorded. |
| shadow | Logs Radix Tree node match ratios without mutating memory state. | Counterfactual match events logged. |
| prod | Enables Radix Tree attention reuse on self-hosted inference nodes. | Active prefix matches and saved prefill tokens logged. |
Current policy
| Memory manager | RadixTree LRU eviction | Dynamic attention node caching. |
Worth knowing before you enable it
- ·Requires high GPU memory allocation for RadixAttention nodes.
What it replaces
- ·Full prefill re-computation on multi-turn chat sessions.
Cluster-wide Radix Tree memory sharing across GPU pools.
- ·Cross-node Radix Tree cache synchronization.
- ·Distributed prefill attention cache routing.