got a 10% (relative) increase in eval scores by simply changing the sampling args to the recommended ones
got a 10% (relative) increase in eval scores by simply changing the sampling args to the recommended ones what are we doing man
got a 10% (relative) increase in eval scores by simply changing the sampling args to the recommended ones what are we doing man
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.