Shared-prefix KV cache
RadixAttention tree-based KV cache reuse across shared prompt prefixes.
problem it solves
Stops redundant prefill compute on recurring system prompts, RAG contexts, and chat histories.
What it does
What it does: Reuses key-value cache tokens across requests sharing common prompt prefixes.
What it watches: Radix tree matches of prefix token sequences in self-hosted model servers.
When it triggers: When incoming request prompt prefix tokens match active KV cache blocks.
The Action: Skips prefill computation for matched prefix tokens, accelerating TTFT.
How it recovers: Unmatched suffix tokens undergo normal prefill processing.
What we need from you
Setup required before this can be enabled
Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.Prefix caching enabled — vLLM `--enable-prefix-caching`, or SGLang, which has RadixAttention on by default.
- Self-hosted vLLM or SGLang deploymentrequired
Requires a model server supporting RadixAttention or prefix caching.
- ACE node agent installedrecommended
Connects control plane telemetry to GPU node cache state.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Your model server computes full prefill for every request. | No prefix_kv_cache stage recorded. |
| shadow | Tracks prefix match potential and logs cache hit ratios without modifying engine state. | Stage with action=would_reuse_prefix logged. |
| prod | Enables prefix KV cache matching on serving nodes, dramatically cutting TTFT. | Prefill tokens saved and TTFT reduction logged. |
Current policy
| Matching structure | Radix tree prefix indexing | Sub-millisecond prefix lookup. |
| Block size | 16 tokens per KV block | Granular memory allocation unit. |
| Eviction policy | LRU with prefix frequency weighting | Retains hot system prompt prefixes. |
Worth knowing before you enable it
- ·Only effective when requests share identical lead token prefixes (system prompt, RAG docs).
- ·Dynamic variables at the start of a prompt break prefix matching for subsequent text.
- ·Requires GPU VRAM allocation for KV cache storage.
What it replaces
- ·Manual context stripping to avoid prefill costs.
- ·Ad-hoc model instance partitioning per system prompt.
- ·Custom prefill compute optimization scripts.
Cross-node KV cache distribution and cache pinning available on enterprise.
- ·Cross-GPU node distributed prefix KV cache synchronization.
- ·Persistent system prompt KV cache pinning.
- ·Advanced VRAM cache partitioning control.