Do extracted circuits actually explain model behavior?
Ablating one task’s discovered circuit hurts another task about as much as ablating that task’s own circuit
Do the circuits we extract to explain a model's behavior actually tell us how it solves a specific task? In new work w/ @nsubramani23 , we find that circuits fail a basic check: ablating one task's circuit hurts another task about as much