Local SLM fallback
Co-located quantized Llama-3-8B emergency backup for zero-downtime cloud provider failover.
problem it solves
Protects application availability when upstream cloud API providers suffer total outages or DNS failures.
What it does
What it does: Reroutes model requests to a co-located local quantized SLM during upstream cloud outages.
What it watches: Cloud API connection health, DNS resolution failures, and 5xx error rate spikes.
When it triggers: When all upstream cloud candidates fail or trip their circuit breakers.
The Action: Dispatches the request to local quantized Llama-3-8B model weights for immediate zero-downtime serving.
How it recovers: Automatically returns to cloud providers once upstream health probes succeed.
What we need from you
- Local GPU / CPU memory provisionedrequired
Co-located node requires VRAM / RAM for quantized 8B model weights.
- Circuit breaker or health probesrecommended
Outage signals trigger fallback routing seamlessly.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Outages return cloud provider errors directly to callers. | No local_slm_fallback stage recorded. |
| shadow | Runs background ping probes against local SLM to verify warmth without using responses. | Stage with action=would_fallback and local_latency logged. |
| prod | Reroutes traffic to local SLM when cloud providers are unreachable. | Served destination indicates local_slm_fallback. |
Current policy
| Local model | Llama-3-8B-Instruct (Q4_K_M) | Co-located GGUF / vLLM local engine. |
| Failover time (ETTR) | <1.2s | Immediate local dispatch on cloud error. |
| VRAM footprint | ~5.5GB VRAM | Kept warm in GPU memory. |
Worth knowing before you enable it
- ·Local 8B models may have different reasoning capabilities than 70B+ cloud models.
- ·Requires local hardware capacity to host quantized model weights.
- ·System prompts must fit within local model context length limits.
What it replaces
- ·Complex multi-cloud fallback routing code.
- ·Custom static response error handlers.
- ·Manual failover runbooks during cloud vendor outages.
Custom local model weights and hybrid fleet configurations available on enterprise.
- ·Deploy fine-tuned local SLM weights.
- ·Multi-GPU local fallback clustering.
- ·Custom hardware acceleration bindings (TRT-LLM / vLLM).