Speculative decoding
Draft-target token sequence generation with EAGLE and small draft models.
problem it solves
Eliminates single-token generation bottlenecks on large self-hosted target models.
What it does
What it does: Accelerates token generation by predicting candidate tokens with a fast draft model.
What it watches: Draft candidate token streams verified in parallel by the target model.
When it triggers: On generation requests targeting large self-hosted LLM endpoints.
The Action: Validates multiple draft tokens per single target model forward pass (1.8x - 2.5x speedup).
How it recovers: Rejected draft tokens are discarded without affecting output mathematical accuracy.
What we need from you
Setup required before this can be enabled
Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.A draft model, or an EAGLE draft head, on the same node class as the target model.
- Target model and compatible draft model/headrequired
Requires draft model (e.g. EAGLE head or Llama-68M draft) co-located on GPU.
- vLLM / SGLang speculative decoding supportrequired
Your model server must support speculative verification kernels.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Target model generates tokens sequentially. | No speculative_decoding stage recorded. |
| shadow | Measures draft acceptance rate on sample requests without altering token pipeline. | Stage with action=would_speculate and acceptance_rate logged. |
| prod | Executes speculative draft verification, boosting tok/s generation speed. | Tokens/sec throughput multiplier and acceptance rate logged. |
Current policy
| Draft mechanism | EAGLE / draft model pairing | Parallel candidate token generation. |
| Lookahead length | K=5 candidate tokens | Optimal verification batch length. |
| Target speedup | 1.8x - 2.5x tok/s | Zero loss in output accuracy or distribution. |
Worth knowing before you enable it
- ·Draft acceptance rates drop on highly creative or non-deterministic prompts.
- ·Requires additional VRAM to hold draft model weights.
- ·Small target models (e.g., 3B) gain less speedup than large 70B+ models.
What it replaces
- ·Traditional single-token autoregressive generation loops.
- ·Manual model distillation for latency reduction.
- ·Complex client-side token streaming hacks.
Adaptive draft model selection and custom EAGLE heads on enterprise.
- ·Custom EAGLE draft head training pipelines.
- ·Dynamic K lookahead length optimization.
- ·Per-workload draft model auto-pairing.