Spectral Lens: looking past loss curves in LLM training
Activation and gradient spectra can expose representation geometry, forecast token efficiency early, and separate real learning gains from throughput gains
New paper: Spectral Lens Loss curves can hide how LLMs actually learn. We show that activation and gradient spectra reveal hidden representation geometry, predict token efficiency early, and distinguish learning gains from throughput gains