This is awesome! This behavior is exactly what we benchmark in
This is awesome! This behavior is exactly what we benchmark in
http://
CodeClash.ai where LMs play against each other in 7 different arenas by writing code. I think there's *so* much more to do in this research direction, and the impacts w
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.