Isabel Silva Corpus

@isabelcorpus.bsky.social

PhD student in Info Sci at Cornell (Tech) isabelsilvacorpus.github.io

🧵Can we reconcile excitement for SAEs with negative results? Our #ICML2026 position paper argues that even if SAEs underperform baselines when acting on knowns (e.g. probing, steering), they're a powerful tool for ~discovering unknowns~ Poster: Tue 2pm, HALL A #1715

Bild

Super excited to share a new #FAccT2026 paper with Nel Escher and Nikola Banovic! We show that algorithm auditing policies in the US are wildly insufficient, as they don't account for how public sector algorithms actually function in practice. dl.acm.org/doi/abs/10.1...

Algorithm Auditing Policies Rest on Flawed Assumptions About Public Sector Systems | Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency

You will be notified whenever a record that you have chosen has been cited.

dl.acm.org

New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246

Bild

I spoke with @kattenbarge.bsky.social for this @wired.com piece about my research into reddit moderators' experiences moderating AI-generated content. Moderators are working hard to keep Reddit "one of the most human spaces left on the internet," but it's a trying and often thankless task.

WIRED@wired.com · 8mo ago

Reddit is considered one of the most human spaces left on the internet, but mods and users are overwhelmed with slop posts in the most popular subreddits. www.wired.com/story/ai-slo...

If you run conjoint experiments, you need to read this. Most conjoints estimate average effects for each attribute. But what if the effect of one attribute depends on the others? This paper has got you covered!

Bild