On-policy distillation with positive-pressure tokens
Using only tokens where the teacher assigns higher probability than the student can still minimize an upper bound on on-policy distillation loss
One nice thing is that you are still minimizing an upper bound of the on-policy distill loss when using only positive pressure tokens (tokens where teacher puts higher probability than student). That said, it's a bit sketchy as per token K