SlimQwen: prune a pretrained MoE rather than train a smaller one
Qwen’s pruning-and-distillation result suggests smaller MoE models can inherit more capability from a large pretrained model than from fresh training
“SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training” This new Qwen paper shows that pruning a pretrained MoE is much better than training the smaller MoE from scratch. All you need to do is prune depth, width