I'm very happy to give a spotlight at the Mechanistic Interpretability Workshop @ ICML on our work: "Validating Causal Abstraction Metrics on Simulated Complex Systems" Which metrics actually tell you if an explanation is valid? We built a benchmark to find out. 1/n
Validating Causal Abstraction Metrics on Simulated Complex Systems
melouxm.github.io