Do extracted circuits actually explain model behavior?

Ablating one task’s discovered circuit hurts another task about as much as ablating that task’s own circuit

Do the circuits we extract to explain a model's behavior actually tell us how it solves a specific task? In new work w/ @nsubramani23 , we find that circuits fail a basic check: ablating one task's circuit hurts another task about as much
Ranked #23 on backlist 2026-05-20 (20 May 2026 UTC) · by (Michael Li) ·

How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.