arXiv — cs.AI preprintsInternational8 October 2026
E$^2$-OPSD: Taming Entropy Overshoot in On-Policy Self-Distillation
This is an official announcement record
Firsthand records what arXiv — cs.AI preprints announced and links to the original. The wording below is theirs, not ours.
arXiv:2610.05048v3 Announce Type: replace-cross Abstract: On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to stu
Read the official announcement
Opens arxiv.org
More from arXiv — cs.AI preprints
- COMPASS: Finding Where Reasoning Lives in Language Models8 October 2026
- Decoupled Multi-Agent Orchestration8 October 2026
- OTel: Open Telco AI Datasets, Benchmarks, and Models8 October 2026
- Comparative review of hybrid forecasting models for short-term prediction of building thermal load8 October 2026
- Dynamical low-rank equilibrium computation for stochastic games between advanced persistent threats and moving target defense8 October 2026
This content is for informational purposes only and is not professional advice. Specifications, prices, plan tiers, and features change frequently and may differ from what is shown here; verify current details on the manufacturer's or company's official page before purchasing. Ratings are based on analysis of published documentation, not independent lab testing.