Prefill–decode disaggregation
Splits compute-heavy prefill nodes from bandwidth-heavy decode nodes.
problem it solves
Stops long prompt prefill compute from blocking short decode token generation.
What it does
What it does: Decouples inference infrastructure into specialized compute (prefill) and memory-bandwidth (decode) GPU pools.
What it watches: Request phase execution state and cluster GPU pool utilization.
When it triggers: On every request handled by disaggregated serving clusters.
The Action: Computes prompt prefill on compute-dense nodes (H100) and streams KV state to decode nodes (H200/L40S).
How it recovers: Falls back to unified nodes if a specialized pool experiences outages.
What we need from you
- At least two node classes registeredrequired
Requires prefill-optimized GPUs (H100) and decode-optimized GPUs/LPUs.
- ACE node agent installedrequired
Manages inter-node KV cache state transfer.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Single nodes handle both prefill and decode phases. | No pd_disaggregation stage recorded. |
| shadow | Tracks phase execution metrics and simulates disaggregated transfer latency. | Stage with action=would_disaggregate logged. |
| prod | Routes prefill and decode phases to specialized GPU pools. | Prefill latency, decode tok/s, and p99 tail reduction logged. |
Current policy
| Prefill pool | H100 / Compute-dense GPUs | Maximizes matrix multiplication FLOPS. |
| Decode pool | H200 / L40S / Groq LPU | Maximizes VRAM memory bandwidth. |
| Tail latency win | Up to 60% p99 reduction | Eliminates queue head-of-line blocking. |
Worth knowing before you enable it
- ·Requires high-speed inter-node interconnects (InfiniBand / RoCE) for KV transfer.
- ·Small prompts gain less benefit than long context RAG and multi-turn requests.
- ·Infrastructure setup requires multi-pool cluster configuration.
What it replaces
- ·Monolithic single-node GPU cluster setups.
- ·Queue head-of-line blocking p99 latency spikes.
- ·Manual inference engine pool partitioning.
Dynamic pool re-balancing and custom interconnect transport on enterprise.
- ·Dynamic node re-allocation between prefill and decode pools.
- ·Custom RDMA / NVLink network transport bindings.
- ·Workload-aware chunked prefill disaggregation.