I find GLM-5.2 currently unusable for hard reasoning tasks. I gave it 11 induction problems from my benchmark (IC…

I find GLM-5.2 currently unusable for hard reasoning tasks. I gave it 11 induction problems from my benchmark (ICML 2026, https:// arxiv.org/abs/2602.18956). - 4 out of the 11 completed, the rest failed; 2 correct - Average time per compl
Ranked #32 on backlist 2026-06-20 (20 Jun 2026 UTC) · by (Serafim Batzoglou) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.