Simple LLM judges break because long-horizon trajectories do not fit into a context window.
Simple LLM judges break because long-horizon trajectories do not fit into a context window.
They either see a narrow slice of the run, or try to ingest a long dense trajectory and miss the evidence in the middle.
Agent Judge gives the eva
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.