Andrew Jesson

@anndvision.bsky.social

does your circuit preserve model behavior ? does removing it disable the task ? does it contain redundant parts ? don ' t know ? then come chat about hypothesis testing for mechanistic interpretability poster 2803 East 4:30-7:30pm at NeurIPS

claudia shi@claudiashi.bsky.social · 2y ago

The circuit hypothesis proposes that LLM capabilities emerge from small subnetworks within the model. But how can we actually test this? 🤔 joint work with @velezbeltran.bsky.social @maggiemakar.bsky.social @anndvision.bsky.social @bleilab.bsky.social Adria @far.ai Achille and Caro

(Shameless) plug for David Blei's lab at Columbia University! People in the lab currently work on a variety of topics, including probabilistic machine learning, Bayesian stats, mechanistic interpretability, causal inference and NLP. Please give us a follow! @bleilab.bsky.social