FP4 training instability comes from Wgrad, not the forward pass
Progressively swapping transformer GEMMs shows MXFP4 full-pipeline training breaks at weight-gradient computation
Full-pipeline FP4 training fails at Wgrad, not the forward pass. This paper isolates MXFP4 instability by progressively replacing FP8 GEMMs in transformer linear layers: - Fprop: forward propagation - Dgrad: activation gradients - Wgra