Elinor

@elinorpd.bsky.social

incoming PhD @ MIT CSAIL // researching LLM societal impacts & alignment previously @ MIT media lab, mila quebec / mcgill i like language and dogs and plants and ultimate frisbee and baking and sunsets. she/her https://elinorp-d.github.io

I spent many hours, days, nights, weekends in cafes, in my office, at home, trying to understand Ran Raz's classical parallel repetition theorem and whether I could prove a quantum version of it. This period of struggle was important for me.

Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/

Bild

We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c

Image of flowers with text overlaid saying: "iNaturalist: We're hiring! Computer vision/machine learning engineer. Full time, remote in the United States."

New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...

I am worried by NLP research culture

In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…

ehudreiter.com

My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision

Ada@adadtur.bsky.social · 2mo ago

Super excited to finally announce my latest research “Would you still call this Dax? Novel Visual References in VLMs and Humans”! We studied how vision-language models (VLMs) adopt new visual concepts and map them to language compared to humans, and found that…

Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...

What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science

statmodeling.stat.columbia.edu

📢 Our paper 🤯🧠 "Brainrot: Deskilling and Addiction are Overlooked AI Risks" has been accepted at the ACM Fairness, Accountability & Transparency (FAccT) conference 2026. The preprint is available: 👉 arxiv.org/abs/2605.03512 TL;DR 🧵 follows 👇 1/5

Brainrot: Deskilling and Addiction are Overlooked AI Risks

The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriat...

arxiv.org

Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :)

Flying out of Boston to Brazil rn means I’m surrounded by poster tubes (ICLR attendees) and blue athletic gear w medals (Boston marathoners). It’s a cool crowd

Congratulations to Jacob Andreas, was named a 2026 Edgerton Award recipient! The award recognizes exceptional teaching, research, and service at MIT! Prof. Andreas co-leads our Language and Thought Mission, and he is a dedicated and creative researcher and educator. news.mit.edu/2026/jacob-a...

Jacob Andreas and Brett McGuire named Edgerton Award winners

MIT associate professors Jacob Andreas and Brett McGuire have been selected as the winners of the 2026 Harold E. Edgerton Faculty Achievement Award for exceptional contributions to teaching, research,...

news.mit.edu

Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.

Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions

To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.

knightcolumbia.org

🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7)

There's been a lot of excitement about pluralistic value alignment 🌈 — AI that reflects the full range of human perspectives But no formal way to benchmark whether we're actually making progress. 🤔 Introducing 𝐎𝐕𝐄𝐑𝐓𝐎𝐍𝐁𝐄𝐍𝐂𝐇. 🎉Accepted to #ICLR2026 1/n 🧵

Bild

Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵

Title, author list, and two figures from the paper. 
Title: The Aftermath of DrawEduMath: Vision Language Models
Underperform with Struggling Students and Misdiagnose Errors
Authors: Li Lucy, Albert Zhang, Nathan Anderson, Ryan Knight, Kyle Lo
Figure 1: On the left is a math problem, where students are asked to draw x < 5/2 on a number line. The right side shows two example student responses that differ in correctness. DrawEduMath pairs each math problem with one student response, and prompts VLMs to answer questions about the student response.
Figure 2: VLMs consistently perform worse on answering DrawEduMath benchmark questions pertaining to erroneous student responses. Performance on non-erroneous student responses is labeled with specific VLMs’ names; that same model’s performance on erroneous student responses is directly below.

Yesterday was my last day at MSR. We recently learned that our roles were eliminated, and with them our little FATE Montreal team. I joined MSR a bit over 7.5 years ago while on active chemotherapy, and being at MSR has overlapped with so much change in my life.