SlimQwen: prune a pretrained MoE rather than train a smaller one

Qwen’s pruning-and-distillation result suggests smaller MoE models can inherit more capability from a large pretrained model than from fresh training

“SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training” This new Qwen paper shows that pruning a pretrained MoE is much better than training the smaller MoE from scratch. All you need to do is prune depth, width
Ranked #13 on backlist 2026-05-17 (17 May 2026 UTC) · by (alphaXiv) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.