FutureSim: replaying the web day by day for continual-learning evals
Agents can be tested on whether they update forecasts as real events unfold rather than on static benchmark snapshots
Continual learning is bottlenecked by realistic evaluations Introducing FutureSim, which replays real-world events in the temporal order they occurred We benchmark frontier agents at updating predictions about how our world evolves, in na