all docs

Feedback distillation ring

Turns the answers your callers accepted into a cheaper model that can serve them.

problem it solves
The router's downgrade threshold is set once, by a benchmark, and then never moves — while your traffic does. A student model that got better last month keeps handling the same small slice, because nobody re-ran the comparison.

What it does

What it does: Collects the responses nothing downstream retried or corrected, uses them as a rolling distillation set for a smaller student model, and moves the router's downgrade threshold as the student's agreement rate rises.

The ring, in three steps

  • ·Collect — accepted responses become training pairs. Acceptance is a signal you send, not a guess we make.
  • ·Distil — the student trains on the slice of your traffic it will actually serve, not on a public benchmark.
  • ·Promote — when agreement on a held-out slice clears the bar, the router sends more traffic to the student. If it falls, the threshold moves back.

Why it compounds: Every skill that saves money does it once. This one raises the ceiling of what llm_router can downgrade, so the saving grows with the traffic rather than staying flat.

What we need from you

  • An acceptance signalrequired

    A thumbs-up, an explicit callback, or a non-retry within a window. Without one there is nothing to learn from, and the ring has no input — this is the hard dependency.

  • llm_router enabledrequired

    The ring moves the router's threshold. With no router there is nothing to move, and the student is trained and never used.

  • Sustained traffic on a stable taskrecommended

    A few thousand accepted responses on work that looks like itself. Traffic that changes shape weekly gives the student a moving target.

What each mode does

ModeEffect on your requestWhat you can see
offNo collection, no training, no threshold movement.No feedback_distillation_ring stage is recorded.
shadowCollection and training run; the router's threshold does not move. The student answers alongside the real model and its agreement rate is recorded — so you can watch it become good enough before it serves anyone.Agreement rate per slice, and the threshold that would have been set.
prodThe threshold moves with the measured agreement rate, in both directions. A student that regresses loses traffic without anyone intervening.Threshold changes with the agreement rate that justified each one.

Current policy

Promotion bar94% agreement on held-outMeasured on a slice the student never trained on.
Evaluation cadencedailyThreshold moves at most once a day, in either direction.
Minimum corpus2,000 accepted responsesBelow this the agreement rate is noise.
Regression responsethreshold reverts immediatelyFalling agreement is acted on faster than rising agreement.

Worth knowing before you enable it

  • ·No acceptance signal means no ring. Non-retry is the cheapest proxy and the one most teams start with.
  • ·It learns your traffic, including its biases. A student distilled from a skewed month serves that month back to you.
  • ·The first threshold move is weeks out, not days — the corpus has to reach the floor first.
  • ·Shadow costs money: the student answers alongside the real model, and both are billed.

What it replaces

  • ·A quarterly re-benchmark of which model handles which query class.
  • ·A fine-tuning pipeline stitched together from exported logs.
  • ·Router thresholds pinned to whatever was true when they were written.

Model ownership and custom promotion policy are available on the enterprise tier.

  • ·Export the distilled student — the weights are yours, servable on your own fleet.
  • ·Your own promotion bar and evaluation slice, including a human-labelled one.
  • ·Per-task rings, so a summarisation student and a classification student are trained apart.
team@acefleet.dev →