Injection guard
Perimeter prompt-injection defense combining fast regex filters and DeBERTa classification.
problem it solves
Blocks instruction overrides, role hijacks, and untrusted retrieval data from hijacking model control flow.
What it does
What it does: Filters malicious prompt injections and jailbreaks before prompts reach model execution.
What it watches: Known attack signatures, role override phrases, system prompt leakage, and vector embeddings.
When it triggers: When input patterns match regex firewall rules or classifier score exceeds threshold.
The Action: Rejects compromised requests with HTTP 400 or strips injection vectors in shadow mode.
How it recovers: Clean inputs pass directly to upstream model providers.
What we need from you
- Regex Firewall Rulesrequired
Deterministic instruction-override patterns run at zero memory overhead.
- Memory for DeBERTa Classifier (Optional)recommended
Learned stage needs ~2.2GB free RAM; falls back to regex-only firewall if constrained.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. All prompts pass through without injection inspection. | No injection_guard stage recorded. |
| shadow | Scores prompts and records potential injection attempts without blocking requests. | Stage with action=would_block and risk_score logged. |
| prod | Refuses requests exceeding injection risk threshold with HTTP 400 Bad Request. | Stage with action=blocked and violation details logged. |
Current policy
| Regex Stage | <0.1ms latency | Catches direct instruction overrides and exfiltration phrases. |
| DeBERTa Classifier | 95.25% accuracy on benchmark | Catches paraphrased and retrieval-borne attacks. |
| Blocking threshold | Score ≥ 0.85 | Configured to minimize false positive blocks. |
Worth knowing before you enable it
- ·Security researchers or red-team prompts will trigger injection defenses.
- ·RAG retrieval context from untrusted websites can trigger indirect injection alerts.
- ·DeBERTa model loading adds ~2.2GB RAM requirement to gateway process.
What it replaces
- ·Custom prompt sanitization middleware.
- ·Third-party LLM firewall API subscriptions.
- ·Ad-hoc system prompt wrapping tricks.
Custom injection signatures and real-time threat feed updates on enterprise.
- ·Custom security policy rule definitions.
- ·Zero-day threat pattern feed updates.
- ·SIEM and security webhook integration.