FP4 training instability comes from Wgrad, not the forward pass

Progressively swapping transformer GEMMs shows MXFP4 full-pipeline training breaks at weight-gradient computation

Full-pipeline FP4 training fails at Wgrad, not the forward pass. This paper isolates MXFP4 instability by progressively replacing FP8 GEMMs in transformer linear layers: - Fprop: forward propagation - Dgrad: activation gradients - Wgra
Ranked #22 on backlist 2026-05-20 (20 May 2026 UTC) Β· by (𝚐π”ͺ𝟾𝚑𝚑𝟾) Β·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.