LLMs have been shown to provide different predictions in clinical tasks when patient race is altered. Can SAEs spot this undue reliance on race? 🧵 Work w/ @byron.bsky.social Link: arxiv.org/abs/2511.00177
Hiba Ahsan
@hibaahsan.bsky.social
PhD student @ Northeastern University, Clinical NLP https://hibaahsan.github.io/ she/her
That’s us! Join Ai2’s Discord AMA on Oct 28, 8 a.m. PT, if you have questions. Paper: arxiv.org/abs/2502.13319
Elucidating Mechanisms of Demographic Bias in LLMs for Healthcare
We know from prior work that LLMs encode social biases, and that this manifests in clinical tasks. In this work we adopt tools from mechanistic interpretability to unveil sociodemographic representati...
arxiv.org
3/ 🏥 A separate team at Northeastern located where certain signals live inside Olmo and made targeted edits that reduced biased clinical predictions. This kind of audit is only possible because Olmo exposes all its components. → buff.ly/HkChr4Q
On the Good Fight podcast w substack.com/@yaschamounk I give a quick but careful primer on how modern AI works. I also chat about our responsibility as machine learning scientists, and what we need to fix to get AI right. Take a listen and reshare - www.persuasion.community/p/david-bau
David Bau on How Artificial Intelligence Works
Yascha Mounk and David Bau delve into the “black box” of AI.
persuasion.community
Who is going to be at #COLM2025? I want to draw your attention to a COLM paper by my student @sfeucht.bsky.social that has totally changed the way I think and teach about LLM representations. The work is worth knowing. And you can meet Sheridan at COLM, Oct 7! bsky.app/profile/sfe...
[📄] Are LLMs mindless token-shifters, or do they build meaningful representations of language? We study how LLMs copy text in-context, and physically separate out two types of induction heads: token heads, which copy literal tokens, and concept heads, which copy word meanings.
I'm searching for some comp/ling experts to provide a precise definition of “slop” as it refers to text (see: corp.oup.com/word-of-the-...) I put together a google form that should take no longer than 10 minutes to complete: forms.gle/oWxsCScW3dJU... If you can help, I'd appreciate your input! 🙏
Oxford Word of the Year 2024 - Oxford University Press
The Oxford Word of the Year 2024 is 'brain rot'. Discover more about the winner, our shortlist, and 20 years of words that reflect the world.
corp.oup.com
LLMs are known to perpetuate social biases in clinical tasks. Can we locate and intervene upon LLM activations that encode patient demographics like gender and race? 🧵 Work w/ @arnabsensharma.bsky.social, @silvioamir.bsky.social, @davidbau.bsky.social, @byron.bsky.social arxiv.org/abs/2502.13319