all docs

Multi-LoRA multiplexing

S-LoRA dynamic adapter weight swapping over a single base model instance.

problem it solves
Prevents duplicating base model VRAM allocations for every customized tenant fine-tune.

What it does

What it does: Dynamically multiplexes hundreds of low-rank LoRA adapters over one shared base model.

What it watches: Incoming request adapter IDs and LoRA weight registries.

When it triggers: On requests specifying a target tenant or task LoRA adapter.

The Action: Loads and applies adapter weights into GPU memory on the fly during inference.

How it recovers: Unrecognized adapter IDs fall back to base model inference.

What we need from you

Setup required before this can be enabled

Needs a model server you run yourself — vLLM, SGLang, or OpenAI-compatible.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.Adapter support — vLLM `--enable-lora`, with your adapters on disk or in a registry ACE can read.
Fleet Registration →
  • Self-hosted base model and LoRA registryrequired

    Requires vLLM / SGLang with S-LoRA dynamic loading enabled.

  • Compatible adapter rank and architecturerequired

    Adapters must target the registered base model.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Requests use fixed static model instances.No multi_lora stage recorded.
shadowVerifies adapter availability and registry state without dynamic weight swapping.Stage with action=would_load_lora logged.
prodMultiplexes requests to the requested LoRA adapter dynamically over shared base model.Active adapter ID and swap latency logged.

Current policy

Multiplexing engineS-LoRA batchingUnified CUDA kernel for multi-adapter execution.
Max active adapters100+ per base modelLimited only by adapter RAM budget.
Swap latency<5ms per adapterAsynchronous weight streaming.

Worth knowing before you enable it

  • ·All fine-tuned adapters must share the identical base model architecture.
  • ·High adapter rank (r > 64) increases VRAM consumption per active tenant.
  • ·Adapter weight files must be accessible over low-latency storage or NVMe.

What it replaces

  • ·Dedicated GPU clusters per fine-tuned customer model.
  • ·Manual model reloading scripts on incoming requests.
  • ·Expensive multi-model deployment manifests.

Automated LoRA weight deployment and isolation vaults on enterprise.

  • ·Automated LoRA registry sync from S3 / HuggingFace.
  • ·Tenant-level adapter access security isolation.
  • ·Dynamic adapter pre-warming based on traffic forecasting.
team@acefleet.dev →