Launching Agentick

Launching Agentick A unified benchmark for training and evaluating general sequential decision-making agents. RL agents, LLMs, VLMs, hybrids, bots, and humans can all be evaluated on: same tasks. same seeds. same score. First result: n
Ranked #30 on backlist 2026-05-12 (12 May 2026 UTC) · by (Roger Creus Castanyer) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.