New AI introspection work with Harvey! Came in skeptical the direct access story would hold but found this series of experiments compelling. (Also, for my fellow 2010s-era psycholinguists: come for the AI introspection, stay for the Brysbaert norms.) arxiv.org/abs/2603.05414
Dissociating Direct Access from Inference in AI Introspection
Introspection is a foundational cognitive ability, but its mechanism is not well understood. Recent work has shown that AI models can introspect. We study their mechanism of introspection, first exten...
arxiv.org
Can large language models *introspect*? In a new paper, @kmahowald.bsky.social and I study the MECHANISM of introspection in big open-source models. tldr: Models detect internal anomalies through DIRECT ACCESS, but don't know what the anomalies are. And they love to guess “apple” 🍎