Dynamic LoRA prefetching
Pre-fetches vLLM/SGLang adapter weights ahead of request execution.
problem it solves
Eliminates weight-swapping cold starts when executing fine-tuned per-tenant models.
What it does
What it does: Asynchronously streams adapter weights from NVMe/S3 to GPU memory before dispatch.
What it watches: Incoming tenant routing metadata and LoRA adapter cache state.
When it triggers: When a request targets a fine-tuned LoRA adapter not yet loaded in GPU RAM.
The Action: Streams adapter weight tensors into S-LoRA memory pools prior to token generation.
How it recovers: Falls back to base model inference if adapter prefetching times out.
What we need from you
Setup required before this can be enabled
Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.Adapter support (`--enable-lora`) with asynchronous adapter streaming.
- 3.Multi-LoRA multiplexing switched on first — prefetching warms the adapters that skill serves.
- Self-hosted multi-LoRA serving clusterrequired
Requires vLLM or SGLang with S-LoRA/Punica adapter streaming.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Adapter weights are loaded synchronously on demand. | No dynamic_lora_prefetch stage recorded. |
| shadow | Tracks adapter cache miss potential without triggering pre-fetches. | Cache miss counterfactual events logged. |
| prod | Triggers predictive async weight pre-fetching into GPU memory pools. | Prefetch latency and adapter hit ratios logged. |
Current policy
| Streaming bandwidth | NVMe PCIe Gen5 / Local SSD | <5ms weight streaming. |
Worth knowing before you enable it
- ·High adapter swapping rates require dedicated PCIe bandwidth.
What it replaces
- ·Synchronous blocking adapter load stalls.
Predictive LoRA pre-warming based on traffic forecasting.
- ·Machine learning traffic predictor for LoRA pre-warming.
- ·Multi-region adapter weight replication.