all docs

Prefill–decode disaggregation

Splits compute-heavy prefill nodes from bandwidth-heavy decode nodes.

problem it solves
Stops long prompt prefill compute from blocking short decode token generation.

What it does

What it does: Decouples inference infrastructure into specialized compute (prefill) and memory-bandwidth (decode) GPU pools.

What it watches: Request phase execution state and cluster GPU pool utilization.

When it triggers: On every request handled by disaggregated serving clusters.

The Action: Computes prompt prefill on compute-dense nodes (H100) and streams KV state to decode nodes (H200/L40S).

How it recovers: Falls back to unified nodes if a specialized pool experiences outages.

What we need from you

  • At least two node classes registeredrequired

    Requires prefill-optimized GPUs (H100) and decode-optimized GPUs/LPUs.

  • ACE node agent installedrequired

    Manages inter-node KV cache state transfer.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Single nodes handle both prefill and decode phases.No pd_disaggregation stage recorded.
shadowTracks phase execution metrics and simulates disaggregated transfer latency.Stage with action=would_disaggregate logged.
prodRoutes prefill and decode phases to specialized GPU pools.Prefill latency, decode tok/s, and p99 tail reduction logged.

Current policy

Prefill poolH100 / Compute-dense GPUsMaximizes matrix multiplication FLOPS.
Decode poolH200 / L40S / Groq LPUMaximizes VRAM memory bandwidth.
Tail latency winUp to 60% p99 reductionEliminates queue head-of-line blocking.

Worth knowing before you enable it

  • ·Requires high-speed inter-node interconnects (InfiniBand / RoCE) for KV transfer.
  • ·Small prompts gain less benefit than long context RAG and multi-turn requests.
  • ·Infrastructure setup requires multi-pool cluster configuration.

What it replaces

  • ·Monolithic single-node GPU cluster setups.
  • ·Queue head-of-line blocking p99 latency spikes.
  • ·Manual inference engine pool partitioning.

Dynamic pool re-balancing and custom interconnect transport on enterprise.

  • ·Dynamic node re-allocation between prefill and decode pools.
  • ·Custom RDMA / NVLink network transport bindings.
  • ·Workload-aware chunked prefill disaggregation.
team@acefleet.dev →