New Method Fixes Hydra Effect Flaw in LLM Circuit Discovery
Five researchers posted "Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models" to arXiv, formally addressing the Hydra effect that has undermined interpretability work since 2023. Their WISE framework and JuntaLearner algorithm estimate a component's true contribution while accounting for compensatory responses elsewhere in the model.
- The Hydra effect, named in a 2023 Google DeepMind paper, sees other components compensate when one is ablated
- WISE borrows the concept of a witness from Pearl and Halpern's causal inference framework
- JuntaLearner ranks components by causal impact in groups rather than individually
- The flaw affects tools like ACDC, EAP and Edge Pruning that score components in isolation
Read next
Hardware