Micah Benson

@micahben.bsky.social

Data Science PhD student @ BU

🧠🤖 The 2026 New England Mechanistic Interpretability (NEMI) Workshop will be Aug. 14 at Boston University! Help spread the word and join the New England mech interp community! Registration and submission info in thread:👇

BU campus and Boston skyline

One of the most common features of AI delusional spirals in our recent study is a belief that the AI is sentient or has a personality. This played a central role in the delusional narratives, and correlated with increased used. Regulators and AI developers should curb this! arxiv.org/abs/2603.16567

Characterizing Delusional Spirals through Human-LLM Chat Logs

As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged in global media and…

arxiv.org

I truly believe the rapid advances in the mech interp subfield have something real to offer AI ethics researchers: A chance to look beyond the HOW of evals to the WHY, a first pass at a technical solution when we see the opportunity, a new avenue for showing failures that prove models are not gods

There's a lot of external pressure on AI ethics to produce solutions instead of critique. As someone who's worked a lot on CSAM, NCII, mental health, and creative harms of AI, if AI developers would have only listened to critiques, we could have avoided all these harms in the first place.