From IcePop to KPop — our team keeps pushing on RL training stability for large MoE models.

From IcePop to KPop — our team keeps pushing on RL training stability for large MoE models. KPop replaces the fixed-ratio mask with an adaptive binary-KL region that matches each token's inherent noise. More robust updates, stable long-ho
Ranked #39 on backlist 2026-05-26 (26 May 2026 UTC) · by (Ant Ling) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.