Why latent prediction can need exponentially less data than token prediction
A sample-complexity theory argues that hidden hierarchical data makes token prediction harder with depth while latent prediction avoids that blowup
"Learn from your own latents, not tokens: A Sample Complexity Theory" This paper explains why data2vec and JEPA can learn with much less data. They showed that when data has hidden hierarchy, token prediction becomes harder as the hierarc