all docs

Injection guard

Perimeter prompt-injection defense combining fast regex filters and DeBERTa classification.

problem it solves
Blocks instruction overrides, role hijacks, and untrusted retrieval data from hijacking model control flow.

What it does

What it does: Filters malicious prompt injections and jailbreaks before prompts reach model execution.

What it watches: Known attack signatures, role override phrases, system prompt leakage, and vector embeddings.

When it triggers: When input patterns match regex firewall rules or classifier score exceeds threshold.

The Action: Rejects compromised requests with HTTP 400 or strips injection vectors in shadow mode.

How it recovers: Clean inputs pass directly to upstream model providers.

What we need from you

  • Regex Firewall Rulesrequired

    Deterministic instruction-override patterns run at zero memory overhead.

  • Memory for DeBERTa Classifier (Optional)recommended

    Learned stage needs ~2.2GB free RAM; falls back to regex-only firewall if constrained.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. All prompts pass through without injection inspection.No injection_guard stage recorded.
shadowScores prompts and records potential injection attempts without blocking requests.Stage with action=would_block and risk_score logged.
prodRefuses requests exceeding injection risk threshold with HTTP 400 Bad Request.Stage with action=blocked and violation details logged.

Current policy

Regex Stage<0.1ms latencyCatches direct instruction overrides and exfiltration phrases.
DeBERTa Classifier95.25% accuracy on benchmarkCatches paraphrased and retrieval-borne attacks.
Blocking thresholdScore ≥ 0.85Configured to minimize false positive blocks.

Worth knowing before you enable it

  • ·Security researchers or red-team prompts will trigger injection defenses.
  • ·RAG retrieval context from untrusted websites can trigger indirect injection alerts.
  • ·DeBERTa model loading adds ~2.2GB RAM requirement to gateway process.

What it replaces

  • ·Custom prompt sanitization middleware.
  • ·Third-party LLM firewall API subscriptions.
  • ·Ad-hoc system prompt wrapping tricks.

Custom injection signatures and real-time threat feed updates on enterprise.

  • ·Custom security policy rule definitions.
  • ·Zero-day threat pattern feed updates.
  • ·SIEM and security webhook integration.
team@acefleet.dev →