Prompt tuning
Measures a model change or a harness migration against a baseline from your own traffic, tunes the prompt to close the gap, and delivers the edit as a PR.
problem it solves
A prompt tuned for one harness and model loses quality when either changes. Without a measured baseline and noise floor nobody can say how big the gap is, or whether an edit closed it.
What it does
Families: AceFleet groups requests that share one prompt — same tenant and agent (`x-ace-agent-name`, or the inferred opening), same `system` and `tools` after variable spans are masked. The spans that vary (dates, ids, per-company blocks) become slots, shown as `{{slot:<id>}}` in the template.
Cases: For the agents you opt in, each decision point is captured redacted (or imported from AI SDK telemetry, OTLP GenAI spans or JSONL): the context up to the step, what the model did, the task's outcome label and its slice (tenant, agent, model, harness). Splits are per task; the held-out set is frozen when a run is created and scored once.
Runs: A run holds arms — one (harness, model, variant) each — and samples every case k times per arm. Model changes replay inside the gateway with recorded tool results; harness migrations run your new harness through `ace eval run`. Every slice gets a paired-bootstrap verdict against the measured noise floor: `better`, `worse`, `within_noise`, or `insufficient` below the slice minimum. Nothing advances on `within_noise`.
The report
- ·Arms × slices verdict matrix with confidence intervals and the noise floor; held-out scores are marked apart from validation scores.
- ·Harness effect (B − A) and model effect (C − B, D − B) attributed separately.
- ·Error rate by class (invalid tool call, failed command, loop, no artifact, refusal, timeout, upstream error), dropout proxies and cost per task.
- ·Quality against cost for every candidate variant, with the frontier marked.
Delivery: The winning variant is a patch anchored to your text plus setting overrides, optionally scoped to one destination model. Open PR turns it into a diff against your own prompt, tool and settings files, each hunk linked to the cases it fixed and broke. Canary 10% through the overlay is optional; once your requests carry the edit the variant is `adopted` and the overlay stops rewriting.
From an agent session: the admin and on-prem MCP servers carry `prompt_tuning_start_run(agent, target_model, budget_usd)` — the family is found by agent name, arm A is the current pair — and `prompt_tuning_proposal(variant_id | run_id)`, which returns the diff with each hunk's fixed and broken cases. Both need the `ace:skills:write` grant on the admin server.
What we need from you
- `x-ace-agent-name` and `x-ace-session` on agent trafficrecommended
The agent name separates families and the session (your task id) keeps every step of a task together, in one split and one canary arm. Without the agent name it is inferred from the conversation opening.
- Outcome labels on at least 30% of casesrequired
POST /v1/tasks/{id}/outcomes with the accepted artifact is the strongest label; OTLP `gen_ai.evaluation.result` scores are ingested too. Promotion needs accepted-artifact or execution labels; inferred outcomes only detect drift.
- A provider key for every target modelrequired
Replay calls are billed to your stored provider key and tagged `source = prompt_tuning_eval`, so they never count as traffic, spend or usage.
- Registered tool schemas for families that call toolsrequired
Register the canonical schemas (exported from Zod or TypeBox) on the family page. Tool-call validity is then checked against them, not against what the SDK happened to send.
- For a harness migration: one `run-case` commandrecommended
It reads `ACE_CASE_FILE`, loads prompts from `ACE_VARIANT_DIR`, uses `ACE_MODEL`, runs writes in dry-run and writes its artifact to `ACE_ARTIFACT_OUT`. Everything in between is yours.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Nothing is captured. Imported cases and on-demand offline runs still work. | No prompt_tuning stage recorded. |
| shadow | Captures redacted cases for the opted-in agents after the response, keeps the baseline and noise floor, and runs evaluations offline. The live request is never changed. | ace_prompt_tuning_overlay_applications_total with mode=shadow; runs, samples and verdicts on the Prompt tuning pages and the prompt_tuning Grafana dashboard. |
| prod | The overlay applies the promoted variant to the whole family (or to N% of sessions under canary:N, each session staying in one arm). A missing anchor makes the variant `stale` and nothing is applied; a request that already carries the edit marks it `adopted`. | Overlay outcome per request (applied, stale, adopted, no_family, no_variant, control) and overlay latency against the 250 ms skill ceiling. |
Current policy
| Samples per case | k = 3 | Set per run; more samples narrow the interval and cost proportionally more. |
| Held-out per reported slice | ≥ 160 labelled cases | Detects a 10-point gap. A smaller slice is reported as `insufficient`, never better or worse. |
| Noise floor | A-vs-A, measured within the last 14 days | Every verdict is against it; a stale floor fails readiness. |
| Baseline before a change | ≥ 7 days | For model- and harness-change triggers. |
| Trigger setting | notify_only | Per family: auto_run, notify_only or off. A person promotes every result. |
Worth knowing before you enable it
- ·A run that fails a readiness check returns `not_ready` and spends $0. The family page lists which check failed and the observed value against its threshold.
- ·A run that reaches its budget stops as `capped` and reports its best variant so far with the share of the search it completed.
- ·The held-out set is scored once per run. A second final score is refused, and no optimiser request ever contains a held-out case.
- ·Protected spans are never edited, and a rule exercised by fewer cases than the slice minimum is listed for review rather than deleted as dead.
- ·Variant text is your content and lives in the variant store; `ace/` refers to variants by id only.
What it replaces
- ·A hand-written eval harness, scorer and statistics code for every model upgrade.
- ·Copy-pasting prompts between harnesses and hoping the new model reads them the same way.
- ·Spreadsheets comparing a handful of runs with no noise floor.
Monthly model qualification and per-company tuning are on the enterprise tier.
- ·Scheduled qualification runs across candidate models with rotating held-out sets.
- ·Per-child-tenant additions tuned on each tenant's own cases.
- ·Runs through an on-prem or local gateway with the same report.