Understanding as compressed representation
The post gives a crisp way to think about why prediction models learn internal structure rather than just raw lookup
A model trained for next-token prediction is forced to build compressed representations of latent structure in text. Ilya Sutskever correctly refers to this phenomenon as understanding. Here, a model trained for next-step sensor prediction,