Multi-LoRA multiplexing
S-LoRA dynamic adapter weight swapping over a single base model instance.
problem it solves
Prevents duplicating base model VRAM allocations for every customized tenant fine-tune.
What it does
What it does: Dynamically multiplexes hundreds of low-rank LoRA adapters over one shared base model.
What it watches: Incoming request adapter IDs and LoRA weight registries.
When it triggers: On requests specifying a target tenant or task LoRA adapter.
The Action: Loads and applies adapter weights into GPU memory on the fly during inference.
How it recovers: Unrecognized adapter IDs fall back to base model inference.
What we need from you
Setup required before this can be enabled
Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.Adapter support — vLLM `--enable-lora`, with your adapters on disk or in a registry ACE can read.
- Self-hosted base model and LoRA registryrequired
Requires vLLM / SGLang with S-LoRA dynamic loading enabled.
- Compatible adapter rank and architecturerequired
Adapters must target the registered base model.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Requests use fixed static model instances. | No multi_lora stage recorded. |
| shadow | Verifies adapter availability and registry state without dynamic weight swapping. | Stage with action=would_load_lora logged. |
| prod | Multiplexes requests to the requested LoRA adapter dynamically over shared base model. | Active adapter ID and swap latency logged. |
Current policy
| Multiplexing engine | S-LoRA batching | Unified CUDA kernel for multi-adapter execution. |
| Max active adapters | 100+ per base model | Limited only by adapter RAM budget. |
| Swap latency | <5ms per adapter | Asynchronous weight streaming. |
Worth knowing before you enable it
- ·All fine-tuned adapters must share the identical base model architecture.
- ·High adapter rank (r > 64) increases VRAM consumption per active tenant.
- ·Adapter weight files must be accessible over low-latency storage or NVMe.
What it replaces
- ·Dedicated GPU clusters per fine-tuned customer model.
- ·Manual model reloading scripts on incoming requests.
- ·Expensive multi-model deployment manifests.
Automated LoRA weight deployment and isolation vaults on enterprise.
- ·Automated LoRA registry sync from S3 / HuggingFace.
- ·Tenant-level adapter access security isolation.
- ·Dynamic adapter pre-warming based on traffic forecasting.