this part is actually very interesting, for the mtp head at t+2 they don't include the kv of the indexer of the p…

this part is actually very interesting, for the mtp head at t+2 they don't include the kv of the indexer of the predicted value at mtp t+1 for efficiency (indexer sharing) AND found that it leads to better results because it avoids training
Ranked #51 on backlist 2026-06-17 (17 Jun 2026 UTC) · by (elie) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.