all docs

Local SLM fallback

Co-located quantized Llama-3-8B emergency backup for zero-downtime cloud provider failover.

problem it solves
Protects application availability when upstream cloud API providers suffer total outages or DNS failures.

What it does

What it does: Reroutes model requests to a co-located local quantized SLM during upstream cloud outages.

What it watches: Cloud API connection health, DNS resolution failures, and 5xx error rate spikes.

When it triggers: When all upstream cloud candidates fail or trip their circuit breakers.

The Action: Dispatches the request to local quantized Llama-3-8B model weights for immediate zero-downtime serving.

How it recovers: Automatically returns to cloud providers once upstream health probes succeed.

What we need from you

  • Local GPU / CPU memory provisionedrequired

    Co-located node requires VRAM / RAM for quantized 8B model weights.

  • Circuit breaker or health probesrecommended

    Outage signals trigger fallback routing seamlessly.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. Outages return cloud provider errors directly to callers.No local_slm_fallback stage recorded.
shadowRuns background ping probes against local SLM to verify warmth without using responses.Stage with action=would_fallback and local_latency logged.
prodReroutes traffic to local SLM when cloud providers are unreachable.Served destination indicates local_slm_fallback.

Current policy

Local modelLlama-3-8B-Instruct (Q4_K_M)Co-located GGUF / vLLM local engine.
Failover time (ETTR)<1.2sImmediate local dispatch on cloud error.
VRAM footprint~5.5GB VRAMKept warm in GPU memory.

Worth knowing before you enable it

  • ·Local 8B models may have different reasoning capabilities than 70B+ cloud models.
  • ·Requires local hardware capacity to host quantized model weights.
  • ·System prompts must fit within local model context length limits.

What it replaces

  • ·Complex multi-cloud fallback routing code.
  • ·Custom static response error handlers.
  • ·Manual failover runbooks during cloud vendor outages.

Custom local model weights and hybrid fleet configurations available on enterprise.

  • ·Deploy fine-tuned local SLM weights.
  • ·Multi-GPU local fallback clustering.
  • ·Custom hardware acceleration bindings (TRT-LLM / vLLM).
team@acefleet.dev →