I find GLM-5.2 currently unusable for hard reasoning tasks. I gave it 11 induction problems from my benchmark (IC…
I find GLM-5.2 currently unusable for hard reasoning tasks. I gave it 11 induction problems from my benchmark (ICML 2026,
https://
arxiv.org/abs/2602.18956).
- 4 out of the 11 completed, the rest failed; 2 correct
- Average time per compl
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.