FutureSim: replaying the web day by day for continual-learning evals

Agents can be tested on whether they update forecasts as real events unfold rather than on static benchmark snapshots

Continual learning is bottlenecked by realistic evaluations Introducing FutureSim, which replays real-world events in the temporal order they occurred We benchmark frontier agents at updating predictions about how our world evolves, in na
Ranked #3 on backlist 2026-05-15 (15 May 2026 UTC) · by (Shashwat Goel) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.