"SWE-bench/ProgramBench are based on publicly-available data, so they're invalid cause the models were trained on…

"SWE-bench/ProgramBench are based on publicly-available data, so they're invalid cause the models were trained on the answers" Nope: 1. Scores are ~0% at first, showing models don't memorize answers. 2. Cheating by post-training on answers
Ranked #77 on backlist 2026-06-04 (04 Jun 2026 UTC) · by (Ofir Press) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.