all docs

Speculative decoding

Draft-target token sequence generation with EAGLE and small draft models.

problem it solves
Eliminates single-token generation bottlenecks on large self-hosted target models.

What it does

What it does: Accelerates token generation by predicting candidate tokens with a fast draft model.

What it watches: Draft candidate token streams verified in parallel by the target model.

When it triggers: On generation requests targeting large self-hosted LLM endpoints.

The Action: Validates multiple draft tokens per single target model forward pass (1.8x - 2.5x speedup).

How it recovers: Rejected draft tokens are discarded without affecting output mathematical accuracy.

What we need from you

Setup required before this can be enabled

Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.A draft model, or an EAGLE draft head, on the same node class as the target model.
Fleet Registration →
  • Target model and compatible draft model/headrequired

    Requires draft model (e.g. EAGLE head or Llama-68M draft) co-located on GPU.

  • vLLM / SGLang speculative decoding supportrequired

    Your model server must support speculative verification kernels.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Target model generates tokens sequentially.No speculative_decoding stage recorded.
shadowMeasures draft acceptance rate on sample requests without altering token pipeline.Stage with action=would_speculate and acceptance_rate logged.
prodExecutes speculative draft verification, boosting tok/s generation speed.Tokens/sec throughput multiplier and acceptance rate logged.

Current policy

Draft mechanismEAGLE / draft model pairingParallel candidate token generation.
Lookahead lengthK=5 candidate tokensOptimal verification batch length.
Target speedup1.8x - 2.5x tok/sZero loss in output accuracy or distribution.

Worth knowing before you enable it

  • ·Draft acceptance rates drop on highly creative or non-deterministic prompts.
  • ·Requires additional VRAM to hold draft model weights.
  • ·Small target models (e.g., 3B) gain less speedup than large 70B+ models.

What it replaces

  • ·Traditional single-token autoregressive generation loops.
  • ·Manual model distillation for latency reduction.
  • ·Complex client-side token streaming hacks.

Adaptive draft model selection and custom EAGLE heads on enterprise.

  • ·Custom EAGLE draft head training pipelines.
  • ·Dynamic K lookahead length optimization.
  • ·Per-workload draft model auto-pairing.
team@acefleet.dev →