Semantic cache
Sub-20ms exact and vector-similarity cache reuse for prompt completions.
problem it solves
Prevents paying full token prices and waiting on model latency for prompts that repeat prior work.
What it does
What it does: Intercepts incoming prompts and serves cached completions when semantic similarity clears a cosine threshold.
What it watches: Vector embeddings of incoming prompt text compared against previously cached request embeddings.
When it triggers: When the normalised prompt matches exactly, or cosine similarity clears the configured threshold (≥0.95 by default).
The Action: Returns the stored completion payload directly in under 20ms with zero upstream model tokens billed.
How it recovers: Cache misses pass through to upstream providers, populating the vector index upon successful response.
What the key covers: The embedded, fuzzy-matched part of the key is the last user turn — the query. Around it, under every policy, the key carries exact boundaries for everything else the answer is a function of: your org and the served model; the **conversation** before the query (a lossy summary of every prior user and assistant turn, so turn five's "and the second one?" at the end of two different conversations is two different questions); the **generation controls** — the system prompt (either shape: a top-level `system`, or `system`/`developer` messages), `tools`, `tool_choice` and `temperature` — so "answer in French" and "answer as JSON" over the same question never share an entry; and any non-text parts (images, audio, files). Then, depending on the deployment's key policy, what the request READ. A single-shot request with none of those is keyed on the query alone.
The conversation summary is a boundary, not a fingerprint. It is deterministic — not an LLM summary, which would cost a model call on the hot path of a cache — and lossy on purpose: NFC-normalised, casefolded, whitespace-collapsed, each turn cut to its first 240 characters. A retyped space or a capital letter does not move the key; a different question does. A repeated conversation hits turn after turn, because once turn one is served from cache turn two's history is byte-identical to the one that wrote it. What it gives up is the cross-conversation hit on a mid-conversation turn — which, measured on a 200-session agent corpus, was the residual 0.3% of served-wrong answers that survived even a caller-declared partition.
Key policy options (ACE_CACHE_KEY_POLICY, set per deployment)
- ·**text** — the default before 2026-09-20. Nothing beyond the always-on boundaries above: the conversation, the generation controls and non-text parts. Right for classification, translation, summarisation, chat and self-contained Q&A: two identical prompts a week apart, under the same instructions, deserve the same answer.
- ·**read_set** — the default. Also covers everything the request READ, from both places it can come from: what the model pulled in during the turn (tool results, MCP results, and the identity of the MCP server that returned them) and what you placed in the prompt for it (`system` and `tools` — the schemas, specs and instructions it was handed). Two turns whose wording is identical but whose context or observations differ stop sharing an answer. Nothing is required of the caller: swap the base URL and this works, because the gateway derives it from the request it already has. A request that read nothing behaves exactly as under text, so switching cannot cold-start a Q&A workload.
- ·**strict** — read_set, plus a refusal: a request that read external state but declared no `x-ace-cache-partition` is not served from the approximate tier at all. Use it when you want the cache to stay out of the way unless a caller has explicitly said which system an answer belongs to — the belt-and-braces setting for agents whose wrong answer is a wrong WRITE. The refused request still runs the shadow lookup, so you see what it would have cost.
What it buys, measured. On a 200-session agent corpus with the shipped cache, embedder and verifier, keying on the prompt alone served a wrong answer on 91.7% of the hits it produced — the cache looked excellent and was almost always wrong, because in an agent session the same sentence means something different each time it is said. `read_set` takes that to **0.0%**. It costs hit rate, deliberately: a narrower key reuses less. Sending an `x-ace-cache-partition` naming the system and its context version raises the hit rate further, but it is an optimisation now, not a requirement — `read_set` is already safe without it.
Choosing between them: one question, and it is not about your industry — would two identical prompts issued a week apart legitimately deserve different answers? No: `text`. Yes, and the answer is read-only: `read_set`. Yes, and the answer causes a write: `strict`.
The multi-turn keying guarantee. A step of an agent loop is never served an earlier or later step's entry. Under `read_set` every tool result, MCP result and `tool`-role message in the request is digested into the key, so step N and step M of one task — same opening prose, same system prompt, same tools, different observations — are two namespaces, on every surface (Anthropic `tool_result` blocks, OpenAI `tool` messages, Responses `function_call_output` items), buffered and streamed, in the exact-match tier and the vector tier alike. A stateful step also hits only inside its own session: a request that read external state, sent with `x-ace-session` and no declared partition, carries the session in its key. So a retry of the same step in the same session is a HIT; a replay of a recorded task under a fresh `x-ace-session` is a MISS on every stateful step — before this, replay k+1 was served replay k's answers step for step, which an evaluation of the skills below the lookup reads as false hits. A request that read nothing keys as if the session were absent, so the org-wide pool for stateless traffic is untouched; a declared `x-ace-cache-partition` outranks the session, so cross-session reuse inside a partition — the measured win for agent verticals — stays. Every verdict says which policy keyed it: `x-ace-cache-key-policy` reads `read_set;ok;read=N` with `;session` appended when the session partition applied, and a HIT carries `x-ace-cache-entry-age-s` beside its timing headers.
What we need from you
- Provider key configured in Provider Vaultrequired
Cache misses require an active upstream provider key to serve initial completions.
- Vector storage backend (Redis / Qdrant)recommended
High-throughput vector store provisioned to hold prompt embeddings and completion payloads.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. Every request passes directly to the upstream model provider. | No semantic_cache stage recorded. |
| shadow | Generates embeddings and evaluates similarity without returning cached payloads to callers. | Stage with action=would_hit recorded on semantic cache matches. |
| prod | Returns cached completions immediately on similarity match, skipping upstream dispatch. | Stage with action=hit recorded; token consumption reported as 0. |
Current policy
| Similarity threshold | ≥0.95 cosine | ACE_CACHE_THRESHOLD. Raised from 0.92 on eval evidence: with the numeric guard on, 0.95 keeps more true hits than pushing the dial alone to 0.98. |
| Cache key policy | read_set | ACE_CACHE_KEY_POLICY — text | read_set | strict. Decides what the cache key covers BESIDES what it always covers (the query, the conversation before it, the system prompt, `tools`, `tool_choice`, `temperature`, and non-text parts): `read_set` adds what the request read through tool and MCP results. Deployment-level — it lives in the gateway's environment, applies to every tenant on that process, and a change is picked up on redeploy. On a dedicated deployment that is one gateway per org, so it is per-customer in practice. |
| Partition threshold | unset (relaxation off) | ACE_CACHE_PARTITION_THRESHOLD. Relaxed cosine floor for lookups inside a caller-declared x-ace-cache-partition, where the namespace already separates "different system" and the dial only has to separate "different question". Deployment-level; read once at gateway startup. |
| Cache TTL | 604,800s (7 days) | Expiry for a stored entry. A read does not extend it. |
| Exact match tier | normalised-prompt map (<1ms) | Checked before the vector lookup: same namespace, same prompt after normalisation, no embedding computed. |
| Embedding model | bge-small-en-v1.5 (local) | Runs in-process via fastembed/ONNX and is shared with the router's classifier. No prompt text leaves the deployment to be embedded. |
Worth knowing before you enable it
- ·Prompts containing dynamic state like timestamps or random seeds degrade similarity match rates.
- ·`temperature` is part of the key, and requests above 0.8 are not SERVED from the cache at all (they still write, and still run the shadow lookup). Two temperatures are two namespaces even below that ceiling: a 0.7 completion is never handed to a caller who asked for 0.
- ·A MISS says why. `x-ace-cache-miss-reason` names the guard that refused the best candidate — `no_candidate`, `below_threshold`, `identifier_mismatch`, `polarity_mismatch`, `negation_mismatch`, `word_order`, `store_error`, or `ceiling_exceeded` / `overloaded` / `loop_stalled` when the lookup did not answer inside the skill latency ceiling (250 ms, per request, not a quota) and the request went upstream rather than wait — and `x-ace-cache-threshold` reports the floor the lookup actually applied (0.98 on an entity-dense prompt, under adaptive thresholds), so a 0.98 similarity beside a MISS reads as a guard working rather than the cache failing.
- ·`identifier_mismatch` compares the identifier tokens of the two prompts as whole tokens, and four shapes make a token an identifier. It carries a digit (`8f2a` ≠ `8e2b`, `T-2026-0919-7` ≠ `T-2026-0919-9`). It carries an underscore, a shape ordinary English does not have (`test_staging` ≠ `test_production` is a different database, and the two score ~0.99 with nothing else here to separate them). It is upper-case and digit-free written after `ticket`, `case`, `incident`, `order`, `id`, `ref`, `sku`, … (`ticket ABX` ≠ `ticket ABY`). Or it is written after a word that names a SLOT rather than a code — `database`, `user`, `branch`, `service`, `instance`, `cluster`, `bucket`, `queue`, `namespace`, `host`, … — where the next token is the name whatever its case, unless it is a function word (`the user in question` names nothing). That last rule is what separates `user alice` from `user bob` and `service payment-gateway` from `service auth-gateway`. A bare hyphen is deliberately not enough: English is full of hyphenated compounds a paraphrase rewrites, so `read-only` still matches `read only`, and a hyphenated machine name is caught where one belongs, after a slot word. A lower-case, letters-only id with no slot word in front of it is still a word to the guard — declare those with `x-ace-cache-partition`.
- ·`polarity_mismatch` and `negation_mismatch` are separate verdicts because they catch different failures. Polarity is two prompts naming OPPOSITE sides of one action while both stay affirmative — `allow` vs `block`, `permit` vs `forbid`, `enable` vs `disable`, `grant` vs `revoke`, `ensure` vs `prevent` — a pair the embedder puts ~0.98 apart because every other word matches. It fires only when each side carries one pole and not the other, so a sentence that names both, the right way to start and stop the cluster, is ordinary phrasing and not a conflict. Negation is a negator present on one side only, or — when both sides negate — one side's negators a proper subset of the other's, the extra negator stacked on a negation they share: is it NOT possible to avoid downtime against is it possible to avoid downtime. Two disjoint negators are left alone, `avoid restarting` and `do not restart` being two ways to say one negative.
- ·`no_candidate` at similarity 0.0 means nothing was reachable, not necessarily nothing was stored. With `ACE_CACHE_ENTITY_PARTITION` on (default off), the gateway namespaces a prompt by its identifier tokens and stamps the digest on the response as `x-ace-cache-entity` — gateway-imposed isolation, distinct from the caller-declared isolation `x-ace-cache-scope` reports. The partition separates exactly the pairs `identifier_mismatch` already refused, so it loses no hit that is served today; what it adds is the ability to tell a cold cache apart from a warm one holding a different entity. Turning it on re-namespaces every stored entry whose prompt carries an identifier token, a partial cold start that is silent by construction.
- ·Provider extensions that shape generation are served; ones that check it are not. `thinking` on the Anthropic surface and `additionalModelRequestFields` / `additionalModelResponseFieldPaths` on the Bedrock Converse surface (where a Bedrock caller puts `thinking`, `top_k`, `reasoning_config`) are part of the key — an entry can only answer a request carrying the identical fields — and read from the cache like any other request. `guardrailConfig` (Bedrock) and `safetySettings` (Gemini) are evaluated by the provider per request, so a request carrying one is `BYPASS` on every call, with the shadow lookup reporting what it would have cost.
- ·The conversation before the query is in the key. A multi-turn workload therefore does not get cross-conversation hits on mid-conversation turns — only a first turn, or a conversation whose prior turns match another's, can hit an entry written elsewhere. That is the trade: the cross-conversation mid-turn hit is precisely the served-wrong answer. Single-shot traffic is unaffected.
- ·Documents pasted into a USER turn are part of the conversation summary only to their first 240 characters. Two conversations that differ only deep inside a pasted document share a key; if that difference changes the answer, declare it with `x-ace-cache-partition` or put it where the key sees it whole — the system prompt.
- ·The key policy and the partition threshold are gateway environment variables, not dashboard toggles. Changing either means redeploying ACE, and the new value applies to every tenant on that process.
- ·Changing the key policy re-namespaces the cache. Entries written under the old policy become unreachable — nothing errors, the hit rate simply drops to zero until it re-warms. Run the new policy in shadow first and compare, rather than switching a busy deployment cold.
- ·Anything volatile in your system prompt or tool list — an interpolated timestamp, a request id, a "current date:" line, an unsorted tool list — gives every request its own cache namespace and the hit rate goes to zero, under every policy. It is a MISS, never a wrong answer, so it fails safely; but it is silent, so the gateway watches for it and logs a warning naming the cause once it sees a run of requests with distinct heads. Same prompt hygiene that provider-side prompt caching needs. (A chat workload is NOT this: its key moves per turn, but its head does not, and the meter reads the head.)
- ·The completion-replay store (the exact-match spend guard on public deployments) uses the same boundaries — conversation, generation controls, media — so nothing is replayable across an instruction or conversation boundary the cache respects.
What it replaces
- ·Custom application-level Redis caching scripts.
- ·Manual embedding similarity comparison code.
- ·Hand-written TTL and eviction management logic.
Custom vector backends and namespace rules are available on the enterprise tier.
- ·Custom embedding models and vector dimensions.
- ·Per-tenant and per-user cache isolation namespaces.
- ·Configurable similarity thresholds per API key.