Knowledge graph
Extracts reusable workflows from successful agent traces, collapsing search loops into deterministic procedural execution.
problem it solves
Without procedural memory, agents repeat costly exploratory trial-and-error steps on every single recurring task.
What it does
What it does: Extracts canonical developer procedures from successful agent execution traces, collapses exploratory search, and prunes superseded steps.
Deterministic scaffolding: Replays verified workflows directly into future runs as structured procedural guidance to eliminate redundant tool calls.
Automatic confidence scoring: Tracks procedure support and success rates with Wilson lower bound confidence intervals to continuously optimize execution paths.
Privacy-safe canonicalisation: Automatically strips volatile tokens, file bodies, and secrets so only safe structural command templates enter the graph.
How it reaches the prompt: in prod, a recalled procedure is rendered as one text block — `[ACE recalled procedure — N prior runs, confidence 0.87]` followed by the numbered steps — and appended to the **last user message**, after the caller's last `cache_control` breakpoint, so the cached prefix is untouched; a later turn of the same session updates that block in place rather than adding another. The block is capped at `ACE_SKG_MAX_TOKENS` (400) and a recall over the cap is `skipped_budget` rather than injected.
What it records: every request the skill looked at carries `kg_*` telemetry in the row's `skill_attrs` — `kg_recall_result` (`hit`, `miss`, `skipped_budget`, `completed`), `kg_matcher`, `kg_match_score`, `kg_procedure_id`, `kg_path_confidence`, `kg_composed_hops`, `kg_injected_tokens`, `kg_cursor_step`, `kg_steps_remaining`, `kg_adherence_delta`, `kg_divergence`, `kg_recall_ms` — and the trace stage `skill_knowledge_graph` (`/v1/execute` with `execution.trace: full`) is the skill's report: `action` (the recall result, or `skipped` with `reason: canary_control` on the control arm, or `failed` with the exception's name when the skill raised and the request went on without it), `injected` (true only when the body that went upstream actually got the block) with `injected_tokens` (0 unless `injected`), `would_inject_tokens` (shadow's counterfactual, under its own name), `block_tokens` against `max_injection_tokens`, the recall itself (`matcher`, `match_score`, `procedure_id`, `path_confidence`, `composed_hops`, `cursor_step`, `steps_remaining`, `adherence_delta`, `divergence`), how the canary arm was decided (`canary_bucket`, `canary_percent`, `canary_keyed_by: session | request`), `session_scoped`, `captured_turns` (what the capture tap buffered for this session) and `duration_ms`. The per-skill scorecard reads the row as `recall_hit_rate`, `adherence_rate`, `avg_injected_tokens` and `divergence_rate`. There is no `x-ace-kg-*` response header; on a provider relay an injection shows on the body-rewrite headers instead — `x-ace-body-rewritten: tail`, `x-ace-rewrite-tokens: +N` and `x-ace-rewritten-by: skill_knowledge_graph` — which are derived from the same stage, so that header names the skill exactly when `injected` is true.
It is a universal skill. It is in the `universal` list `GET /api/v1/skills` returns and accepts `x-ace-skills: skill_knowledge_graph=off` like any other — and it is what `x-ace-skills: *=off` switches off along with the rest. A control arm that names skills individually must name this one too, or it keeps injecting.
What we need from you
- Multi-turn agent sessionsrequired
Operates on multi-turn tool calling and command execution traces. Capture and recall are keyed by `x-ace-session` (or `session_id` in the body); a request without one is not buffered and recalls nothing.
- Task outcomerecommended
Send `x-ace-task-outcome: success` (or `failure`; `x-ace-outcome` / `x-ace-task-success` are accepted spellings) on a task's final call so the trace is scored. Idle sessions are flushed and mined by a background sweep every minute, after 10 minutes without a turn.
- Clean verification signalsrecommended
Identifies procedure completion via verify steps (e.g. pytest, typechecks, builds).
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted: no capture, no recall. Agents execute full exploratory search without procedural recall. | No skill_knowledge_graph stage recorded; no kg_* fields on the row. |
| shadow | Captures turns and mines procedures, and runs the recall on every request — but injects nothing; the prompt goes upstream as sent. | The row's kg_* fields carry the recall (`kg_recall_result: hit`, the procedure id) and `kg_injected_tokens` keeps the size of the block that WOULD have been injected — deliberately, because the dashboards size the skill on that column. The trace stage separates the two: `injected: false`, `injected_tokens: 0`, and the same would-be figure as `would_inject_tokens`. The row and the stage disagree in shadow by design; the stage is the one that says what went upstream, which is nothing. |
| prod | Injects the recalled procedure block into the last user message after the caller's cache breakpoint, tracks the cursor through the procedure across the session, and records adherence and divergence. | `kg_recall_result: hit` with `kg_injected_tokens` > 0 when the block went in; the trace stage carries `action: hit`, `injected: true`, `injected_tokens` equal to `block_tokens` (under `max_injection_tokens`), the recall fields and `captured_turns`. A lever write that did not take reads `injected: false` with `injected_tokens: 0` even on a hit. A block over the cap is `action: skipped_budget` with `block_tokens` and `max_injection_tokens` beside it. On a provider relay the same hit shows as `x-ace-body-rewritten: tail`, `x-ace-rewrite-tokens: +N`, `x-ace-rewritten-by: skill_knowledge_graph`. |
| canary_control | The key has the skill on under `canary:N` and this session hashed to control: nothing ran — no capture, no recall, the prompt goes upstream as sent. | A stage rather than an absence, so control cannot be mistaken for off: `action: skipped`, `reason: canary_control`, `mode: canary_control`, with `canary_bucket: control`, `canary_percent: N`, `canary_keyed_by: session` (or `request` when there was no session) and `session_scoped`. No kg_* recall fields on it; `x-ace-skill-modes` reports `skill_knowledge_graph=canary_control`. |
Current policy
| Exploration collapse | Enabled | Prunes read-only exploratory steps and hoists discovered targets. |
| Confidence model | Wilson 95% lower bound | Requires statistical support before promoting procedures to recallable state. |
| Token budget cap | 400 tokens (`ACE_SKG_MAX_TOKENS`) | Bounds injected procedural guidance to prevent context bloating. |
Worth knowing before you enable it
- ·The default store is in-memory per process; a restart empties it, and on a multi-replica deployment each replica mines and recalls its own graph (a Postgres store is the durable, shared backend).
- ·Under `canary:N` this is the one skill whose bucket is resolved per session rather than per request — a session is treatment or control for its whole run, so a procedure is never injected on step 3 and withheld on step 4; `x-ace-skill-modes` reports `canary_treatment` or `canary_control` accordingly, and the trace stage says why (`canary_bucket`, `canary_percent`, and `canary_keyed_by: request` when a request without a session fell back to per-request hashing).
- ·The injected block changes the last user message. A caller that hashes or replays the exact bytes of its final turn, or asserts the prompt went upstream verbatim, must run this skill `off` (or `*=off`) on that traffic — a docs off-list that predates this skill will not have named it.
- ·Tasks with non-deterministic or rapidly changing environments require fresh exploration.
- ·Only procedures that terminate in a verified passing step are saved to the graph.
What it replaces
- ·Hardcoded runbooks and brittle bash script automations.
- ·Redundant LLM exploratory loops on recurring agent tasks.
Cross-organization procedural memory graphs and custom knowledge extraction.
- ·Federated procedural memory sharing across teams.
- ·Custom argument allowlists and proprietary domain tool normalizers.
- ·Dedicated high-throughput graph database backends.