all docs

Radix tree attention cache

SGLang Radix Tree attention memory reuse for complex multi-turn prompt contexts.

problem it solves
Prevents redundant prefill attention calculation across branching multi-turn chat sessions and long-context prompts.

What it does

What it does: Reuses Radix Tree attention nodes across branching multi-turn chat prompts.

What it watches: Prefix tree matches on self-hosted SGLang serving clusters.

When it triggers: When incoming prompt tokens share prefix nodes with active Radix Tree branches.

The Action: Skips prefill computation for matching prefix branches, drastically reducing TTFT.

How it recovers: Non-matching branch extensions undergo standard prefill processing.

What we need from you

Setup required before this can be enabled

Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.SGLang, or vLLM built with RadixTree memory management and radix attention left enabled.
Fleet Registration →
  • Self-hosted SGLang or vLLM deploymentrequired

    Requires a model server with a RadixTree memory manager.

What each mode does

ModeEffect on your requestWhat you can see
offYour model server re-computes full prefill attention on every request.No radix_cache_attention stage recorded.
shadowLogs Radix Tree node match ratios without mutating memory state.Counterfactual match events logged.
prodEnables Radix Tree attention reuse on self-hosted inference nodes.Active prefix matches and saved prefill tokens logged.

Current policy

Memory managerRadixTree LRU evictionDynamic attention node caching.

Worth knowing before you enable it

  • ·Requires high GPU memory allocation for RadixAttention nodes.

What it replaces

  • ·Full prefill re-computation on multi-turn chat sessions.

Cluster-wide Radix Tree memory sharing across GPU pools.

  • ·Cross-node Radix Tree cache synchronization.
  • ·Distributed prefill attention cache routing.
team@acefleet.dev →