all docs

Teacher–student distillation

Offloads frontier API workloads to fine-tuned compact student models.

problem it solves
Eliminates high recurring frontier API billing for domain-specific inference tasks.

What it does

What it does: Trains compact student models from frontier teacher outputs and routes traffic once accuracy parity is met.

What it watches: Logged teacher prompt-completion pairs and evaluation accuracy scores.

When it triggers: When student model accuracy evaluation clears configured benchmark bar (e.g., ≥95%).

The Action: Switches production traffic from paid frontier APIs to self-hosted student model endpoints.

How it recovers: Low-confidence student predictions fall back to teacher API endpoints.

What we need from you

Setup required before this can be enabled

Needs somewhere to serve the student model — usually a model server you run yourself.

  1. 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
  2. 2.Request logging retained at least 30 days — the student is trained from your own traffic.
  3. 3.An endpoint to serve the student once it clears your accuracy bar.
Fleet Registration →
  • Request logging retained for 30+ daysrequired

    Provides dataset for student model fine-tuning.

  • Self-hosted student endpointrequired

    Serving capacity (e.g. 8B/14B model) to run distilled student model.

What each mode does

ModeEffect on your requestWhat you can see
offNot consulted. All requests continue calling frontier teacher APIs.No distillation stage recorded.
shadowLogs teacher outputs and scores parallel student predictions without serving student outputs.Stage with action=would_distill and parity_score logged.
prodRoutes qualified domain requests to self-hosted student model.Frontier API token savings and student accuracy logged.

Current policy

Student model class8B / 14B compact LLMHigh throughput, low self-hosted cost.
Parity threshold≥95% accuracy matchVerified against held-out test evaluation set.
Data extractionAutomated log pipelinePrivacy-preserving prompt log collector.

Worth knowing before you enable it

  • ·Student models require initial training dataset collection from live traffic.
  • ·Broad open-ended reasoning tasks distill less effectively than structured domain tasks.
  • ·Periodic re-distillation is required as domain requirements evolve.

What it replaces

  • ·Manual dataset export and offline fine-tuning scripts.
  • ·Hand-crafted eval benchmark suites.
  • ·High monthly frontier API expenditure.

Automated distillation pipelines and evaluation suites on enterprise.

  • ·Continuous automated student model retraining.
  • ·Automated synthetic dataset expansion.
  • ·Enterprise model registry and deployment automation.
team@acefleet.dev →