all docs

Shared-prefix KV cache

RadixAttention tree-based KV cache reuse across shared prompt prefixes.

problem it solves
Stops redundant prefill compute on recurring system prompts, RAG contexts, and chat histories.

What it does

What it does: Reuses key-value cache tokens across requests sharing common prompt prefixes.

What it watches: Radix tree matches of prefix token sequences in self-hosted model servers.

When it triggers: When incoming request prompt prefix tokens match active KV cache blocks.

The Action: Skips prefill computation for matched prefix tokens, accelerating TTFT.

How it recovers: Unmatched suffix tokens undergo normal prefill processing.

What we need from you

Setup required before this can be enabled

Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.Prefix caching enabled — vLLM `--enable-prefix-caching`, or SGLang, which has RadixAttention on by default.
Fleet Registration →
  • Self-hosted vLLM or SGLang deploymentrequired

    Requires a model server supporting RadixAttention or prefix caching.

  • ACE node agent installedrecommended

    Connects control plane telemetry to GPU node cache state.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Your model server computes full prefill for every request.No prefix_kv_cache stage recorded.
shadowTracks prefix match potential and logs cache hit ratios without modifying engine state.Stage with action=would_reuse_prefix logged.
prodEnables prefix KV cache matching on serving nodes, dramatically cutting TTFT.Prefill tokens saved and TTFT reduction logged.

Current policy

Matching structureRadix tree prefix indexingSub-millisecond prefix lookup.
Block size16 tokens per KV blockGranular memory allocation unit.
Eviction policyLRU with prefix frequency weightingRetains hot system prompt prefixes.

Worth knowing before you enable it

  • ·Only effective when requests share identical lead token prefixes (system prompt, RAG docs).
  • ·Dynamic variables at the start of a prompt break prefix matching for subsequent text.
  • ·Requires GPU VRAM allocation for KV cache storage.

What it replaces

  • ·Manual context stripping to avoid prefill costs.
  • ·Ad-hoc model instance partitioning per system prompt.
  • ·Custom prefill compute optimization scripts.

Cross-node KV cache distribution and cache pinning available on enterprise.

  • ·Cross-GPU node distributed prefix KV cache synchronization.
  • ·Persistent system prompt KV cache pinning.
  • ·Advanced VRAM cache partitioning control.
team@acefleet.dev →