Teacher–student distillation
Offloads frontier API workloads to fine-tuned compact student models.
problem it solves
Eliminates high recurring frontier API billing for domain-specific inference tasks.
What it does
What it does: Trains compact student models from frontier teacher outputs and routes traffic once accuracy parity is met.
What it watches: Logged teacher prompt-completion pairs and evaluation accuracy scores.
When it triggers: When student model accuracy evaluation clears configured benchmark bar (e.g., ≥95%).
The Action: Switches production traffic from paid frontier APIs to self-hosted student model endpoints.
How it recovers: Low-confidence student predictions fall back to teacher API endpoints.
What we need from you
Setup required before this can be enabled
Needs somewhere to serve the student model — usually a model server you run yourself.
- 1.Add your model server on Fleet Registration — name it, paste the endpoint, pick vLLM / SGLang / OpenAI-compatible. ACE probes it from there to see what it supports.
- 2.Request logging retained at least 30 days — the student is trained from your own traffic.
- 3.An endpoint to serve the student once it clears your accuracy bar.
- Request logging retained for 30+ daysrequired
Provides dataset for student model fine-tuning.
- Self-hosted student endpointrequired
Serving capacity (e.g. 8B/14B model) to run distilled student model.
What each mode does
| Mode | Effect on your request | What you can see |
|---|---|---|
| off | Not consulted. All requests continue calling frontier teacher APIs. | No distillation stage recorded. |
| shadow | Logs teacher outputs and scores parallel student predictions without serving student outputs. | Stage with action=would_distill and parity_score logged. |
| prod | Routes qualified domain requests to self-hosted student model. | Frontier API token savings and student accuracy logged. |
Current policy
| Student model class | 8B / 14B compact LLM | High throughput, low self-hosted cost. |
| Parity threshold | ≥95% accuracy match | Verified against held-out test evaluation set. |
| Data extraction | Automated log pipeline | Privacy-preserving prompt log collector. |
Worth knowing before you enable it
- ·Student models require initial training dataset collection from live traffic.
- ·Broad open-ended reasoning tasks distill less effectively than structured domain tasks.
- ·Periodic re-distillation is required as domain requirements evolve.
What it replaces
- ·Manual dataset export and offline fine-tuning scripts.
- ·Hand-crafted eval benchmark suites.
- ·High monthly frontier API expenditure.
Automated distillation pipelines and evaluation suites on enterprise.
- ·Continuous automated student model retraining.
- ·Automated synthetic dataset expansion.
- ·Enterprise model registry and deployment automation.