@BowenWangNLP
@BowenWangNLP et al. dropped 32,122 verifiable rlvr tasks for training cua agents which is about 87x of osworld tasks. large enough to experiment some cua rl scaling
@BowenWangNLP et al. dropped 32,122 verifiable rlvr tasks for training cua agents which is about 87x of osworld tasks. large enough to experiment some cua rl scaling
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.