On-policy distillation with positive-pressure tokens

Using only tokens where the teacher assigns higher probability than the student can still minimize an upper bound on on-policy distillation loss

One nice thing is that you are still minimizing an upper bound of the on-policy distill loss when using only positive pressure tokens (tokens where teacher puts higher probability than student). That said, it's a bit sketchy as per token K
Ranked #18 on backlist 2026-06-13 (13 Jun 2026 UTC) · by (Rishabh Agarwal) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.