Prompt tuning

Measures a model change or a harness migration against a baseline from your own traffic, tunes the prompt to close the gap, and delivers the edit as a PR.

problem it solves
A prompt tuned for one harness and model loses quality when either changes. Without a measured baseline and noise floor nobody can say how big the gap is, or whether an edit closed it.

What it does

Families: AceFleet groups requests that share one prompt — same tenant and agent (`x-ace-agent-name`, or the inferred opening), same `system` and `tools` after variable spans are masked. The spans that vary (dates, ids, per-company blocks) become slots, shown as `{{slot:<id>}}` in the template.

Cases: For the agents you opt in, each decision point is captured redacted (or imported from AI SDK telemetry, OTLP GenAI spans or JSONL): the context up to the step, what the model did, the task's outcome label and its slice (tenant, agent, model, harness). Splits are per task; the held-out set is frozen when a run is created and scored once.

Runs: A run holds arms — one (harness, model, variant) each — and samples every case k times per arm. Model changes replay inside the gateway with recorded tool results; harness migrations run your new harness through `ace eval run`. Every slice gets a paired-bootstrap verdict against the measured noise floor: `better`, `worse`, `within_noise`, or `insufficient` below the slice minimum. Nothing advances on `within_noise`.

The report

  • ·Arms × slices verdict matrix with confidence intervals and the noise floor; held-out scores are marked apart from validation scores.
  • ·Harness effect (B − A) and model effect (C − B, D − B) attributed separately.
  • ·Error rate by class (invalid tool call, failed command, loop, no artifact, refusal, timeout, upstream error), dropout proxies and cost per task.
  • ·Quality against cost for every candidate variant, with the frontier marked.

Delivery: The winning variant is a patch anchored to your text plus setting overrides, optionally scoped to one destination model. Open PR turns it into a diff against your own prompt, tool and settings files, each hunk linked to the cases it fixed and broke. Canary 10% through the overlay is optional; once your requests carry the edit the variant is `adopted` and the overlay stops rewriting.

From an agent session: the admin and on-prem MCP servers carry `prompt_tuning_start_run(agent, target_model, budget_usd)` — the family is found by agent name, arm A is the current pair — and `prompt_tuning_proposal(variant_id | run_id)`, which returns the diff with each hunk's fixed and broken cases. Both need the `ace:skills:write` grant on the admin server.

What we need from you

  • `x-ace-agent-name` and `x-ace-session` on agent trafficrecommended

    The agent name separates families and the session (your task id) keeps every step of a task together, in one split and one canary arm. Without the agent name it is inferred from the conversation opening.

  • Outcome labels on at least 30% of casesrequired

    POST /v1/tasks/{id}/outcomes with the accepted artifact is the strongest label; OTLP `gen_ai.evaluation.result` scores are ingested too. Promotion needs accepted-artifact or execution labels; inferred outcomes only detect drift.

  • A provider key for every target modelrequired

    Replay calls are billed to your stored provider key and tagged `source = prompt_tuning_eval`, so they never count as traffic, spend or usage.

  • Registered tool schemas for families that call toolsrequired

    Register the canonical schemas (exported from Zod or TypeBox) on the family page. Tool-call validity is then checked against them, not against what the SDK happened to send.

  • For a harness migration: one `run-case` commandrecommended

    It reads `ACE_CASE_FILE`, loads prompts from `ACE_VARIANT_DIR`, uses `ACE_MODEL`, runs writes in dry-run and writes its artifact to `ACE_ARTIFACT_OUT`. Everything in between is yours.

What each mode does

ModeEffect on your requestWhat you can see
off
Nothing is captured. Imported cases and on-demand offline runs still work.
No prompt_tuning stage recorded.
shadow
Captures redacted cases for the opted-in agents after the response, keeps the baseline and noise floor, and runs evaluations offline. The live request is never changed.
ace_prompt_tuning_overlay_applications_total with mode=shadow; runs, samples and verdicts on the Prompt tuning pages and the prompt_tuning Grafana dashboard.
prod
The overlay applies the promoted variant to the whole family (or to N% of sessions under canary:N, each session staying in one arm). A missing anchor makes the variant `stale` and nothing is applied; a request that already carries the edit marks it `adopted`.
Overlay outcome per request (applied, stale, adopted, no_family, no_variant, control) and overlay latency against the 250 ms skill ceiling.

Current policy

Samples per casek = 3
Set per run; more samples narrow the interval and cost proportionally more.
Held-out per reported slice≥ 160 labelled cases
Detects a 10-point gap. A smaller slice is reported as `insufficient`, never better or worse.
Noise floorA-vs-A, measured within the last 14 days
Every verdict is against it; a stale floor fails readiness.
Baseline before a change≥ 7 days
For model- and harness-change triggers.
Trigger settingnotify_only
Per family: auto_run, notify_only or off. A person promotes every result.

Worth knowing before you enable it

  • ·A run that fails a readiness check returns `not_ready` and spends $0. The family page lists which check failed and the observed value against its threshold.
  • ·A run that reaches its budget stops as `capped` and reports its best variant so far with the share of the search it completed.
  • ·The held-out set is scored once per run. A second final score is refused, and no optimiser request ever contains a held-out case.
  • ·Protected spans are never edited, and a rule exercised by fewer cases than the slice minimum is listed for review rather than deleted as dead.
  • ·Variant text is your content and lives in the variant store; `ace/` refers to variants by id only.

What it replaces

  • ·A hand-written eval harness, scorer and statistics code for every model upgrade.
  • ·Copy-pasting prompts between harnesses and hoping the new model reads them the same way.
  • ·Spreadsheets comparing a handful of runs with no noise floor.

Monthly model qualification and per-company tuning are on the enterprise tier.

  • ·Scheduled qualification runs across candidate models with rotating held-out sets.
  • ·Per-child-tenant additions tuned on each tenant's own cases.
  • ·Runs through an on-prem or local gateway with the same report.
team@acefleet.dev →