"SWE-bench/ProgramBench are based on publicly-available data, so they're invalid cause the models were trained on…
"SWE-bench/ProgramBench are based on publicly-available data, so they're invalid cause the models were trained on the answers"
Nope:
1. Scores are ~0% at first, showing models don't memorize answers.
2. Cheating by post-training on answers
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.