all docs

Dynamic LoRA prefetching

Pre-fetches vLLM/SGLang adapter weights ahead of request execution.

problem it solves
Eliminates weight-swapping cold starts when executing fine-tuned per-tenant models.

What it does

What it does: Asynchronously streams adapter weights from NVMe/S3 to GPU memory before dispatch.

What it watches: Incoming tenant routing metadata and LoRA adapter cache state.

When it triggers: When a request targets a fine-tuned LoRA adapter not yet loaded in GPU RAM.

The Action: Streams adapter weight tensors into S-LoRA memory pools prior to token generation.

How it recovers: Falls back to base model inference if adapter prefetching times out.

What we need from you

Setup required before this can be enabled

Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.Adapter support (`--enable-lora`) with asynchronous adapter streaming.
  3. 3.Multi-LoRA multiplexing switched on first — prefetching warms the adapters that skill serves.
Fleet Registration →
  • Self-hosted multi-LoRA serving clusterrequired

    Requires vLLM or SGLang with S-LoRA/Punica adapter streaming.

What each mode does

ModeEffect on your requestWhat you can see
offAdapter weights are loaded synchronously on demand.No dynamic_lora_prefetch stage recorded.
shadowTracks adapter cache miss potential without triggering pre-fetches.Cache miss counterfactual events logged.
prodTriggers predictive async weight pre-fetching into GPU memory pools.Prefetch latency and adapter hit ratios logged.

Current policy

Streaming bandwidthNVMe PCIe Gen5 / Local SSD<5ms weight streaming.

Worth knowing before you enable it

  • ·High adapter swapping rates require dedicated PCIe bandwidth.

What it replaces

  • ·Synchronous blocking adapter load stalls.

Predictive LoRA pre-warming based on traffic forecasting.

  • ·Machine learning traffic predictor for LoRA pre-warming.
  • ·Multi-region adapter weight replication.
team@acefleet.dev →