I spent many hours, days, nights, weekends in cafes, in my office, at home, trying to understand Ran Raz's classical parallel repetition theorem and whether I could prove a quantum version of it. This period of struggle was important for me.
Elinor
@elinorpd.bsky.social
incoming PhD @ MIT CSAIL // researching LLM societal impacts & alignment previously @ MIT media lab, mila quebec / mcgill i like language and dogs and plants and ultimate frisbee and baking and sunsets. she/her https://elinorp-d.github.io
Pluralistic alignment is thriving as a research agenda yet failing at its goal: making the AI systems people actually use more pluralistic🌈 🚨New position paper: we argue adoption in deployed models should be the fields main goal & we provide a roadmap of how to get there 🧵1/
With my ACL PC abilities, I wanted to characterize who these folks are, whether they're actually "outsiders", whether their papers are much less likely to be accepted, is this growth just LLM slop? The whole analsis is here medium.com/@jurgens_245...
Is the ACL Rolling Review actually broken?
ARR is broken! ARR is being overwhelmed with papers from outsiders! LLM-generated slop is ruining our peer review! There are not enough…
medium.com
GreenEarth is creating open source AI-driven recommender infrastructure for BlueSky. Type a prompt, see your feed change. We are here for the users, the builders, the dreamers. Join us. greenearthsocial.substack.com/p/introducin...
Introducing GreenEarth
We're building advanced open source algorithms for social media
greenearthsocial.substack.com
We're looking for a CV/ML Engineer to help us improve the machine learning systems that power iNaturalist's species identification and geographic range modeling. If you're excited to help build tools that help millions of people engage with nature, we'd love to hear from you! Apply: buff.ly/YZqaW6c
New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...
I am worried by NLP research culture
In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…
ehudreiter.com
excited to share that i'll be pursuing my phd in computer science at @csail.mit.edu starting this fall 🥳 🎓 i'm so grateful to be coadvised by the literal dream team: Jacob Andreas, @mbakker.bsky.social, and @mitchellgordon.bsky.social 🙌
My first last-author paper is out! If you saw this dog below and someone showed you the second image, would you consider them the same word/concept? (more examples in Ada's thread) We study if VLMs agree with humans on this and revisit old questions around shape vs. texture bias in vision
Super excited to finally announce my latest research “Would you still call this Dax? Novel Visual References in VLMs and Humans”! We studied how vision-language models (VLMs) adopt new visual concepts and map them to language compared to humans, and found that…
Thoughts on metascientific consequences of AI-generated slides & ideas diluting the impression that speakers are commited to what they present. Science runs on personal attachment more than we admit. If it were a cake mix, how wouldn we add back an egg? statmodeling.stat.columbia.edu/2026/05/28/w...
What if scientists really were dispassionate observers, communicating ideas without irrational commitment? Look here, says AI. | Statistical Modeling, Causal Inference, and Social Science
statmodeling.stat.columbia.edu
📢 Our paper 🤯🧠 "Brainrot: Deskilling and Addiction are Overlooked AI Risks" has been accepted at the ACM Fairness, Accountability & Transparency (FAccT) conference 2026. The preprint is available: 👉 arxiv.org/abs/2605.03512 TL;DR 🧵 follows 👇 1/5
Brainrot: Deskilling and Addiction are Overlooked AI Risks
The scope of AI safety and alignment work in generative artificial intelligence (GenAI) has so far mostly been limited to harms related to: (a) discrimination and hate speech, (b) harmful/inappropriat...
arxiv.org
Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026
LLMs have been widely reported as left-wing biased. The finding has shaped policy and debate — with Trump banning "Woke AI". Our new paper challenges this story. It's not that the models are biased. It's that they think the auditor is. 🧵 w/ michelleschimmel.bsky.socialarxiv.org/pdf/2604.27633
Our position paper, "AI Welfare Is Bullshit" just got accepted to ICML 2026 @icmlconf.bsky.social! The AI welfare agenda has already begun to attract institutional investment at organizations such an @anthropic.com. We argue that this idea is essentially Frankfurtian Bullshit.
Designating that most uses of the term "goblin" and "gremlin" are not legitimate while designating most uses of the word "frog" as legitimate is the imposition of a set of values on a technology. Why do a small number of people in Silicon Valley get to decide whether goblins are inappropriate?
most uses of frog turned out to be legitimate.
Patrick Sui and I are hosting an #ICLR2026 social for anyone with background/interest in the humanities! Room 210, 12-1:30pm on Friday 24 April!! Humanities-adjacent, humanities-curious, everyone is welcome! Should be a fun group! :)
Flying out of Boston to Brazil rn means I’m surrounded by poster tubes (ICLR attendees) and blue athletic gear w medals (Boston marathoners). It’s a cool crowd
I'll be presenting OvertonBench at #ICLR2026 in Rio later this week! 📍Sat, Apr 25, 10:30am in Pavilion 4 (#4109) Please DM me if you'd like to chat about pluralistic / value alignment, societal impacts, epistemology, fairness, evals, etc
There's been a lot of excitement about pluralistic value alignment 🌈 — AI that reflects the full range of human perspectives But no formal way to benchmark whether we're actually making progress. 🤔 Introducing 𝐎𝐕𝐄𝐑𝐓𝐎𝐍𝐁𝐄𝐍𝐂𝐇. 🎉Accepted to #ICLR2026 1/n 🧵
Congratulations to Jacob Andreas, was named a 2026 Edgerton Award recipient! The award recognizes exceptional teaching, research, and service at MIT! Prof. Andreas co-leads our Language and Thought Mission, and he is a dedicated and creative researcher and educator. news.mit.edu/2026/jacob-a...
Jacob Andreas and Brett McGuire named Edgerton Award winners
MIT associate professors Jacob Andreas and Brett McGuire have been selected as the winners of the 2026 Harold E. Edgerton Faculty Achievement Award for exceptional contributions to teaching, research,...
news.mit.edu
My Hoya Rebecca put out the cutest teeny lil flowers today and I’m obsessed
Tech industry mottos have a mixed track record. But we should hold idealists to their ideals. And we should celebrate when they come through. The Mythos non-release is a remarkable moment of conviction. Thoughts: davidbau.com/archives/20... Bravo to Anthropic's "race the top".
Democracy isn't a rulebook. It runs on daily interactions where people comply with norms and hold each other accountable. AI agents are about to join that system. We need to build them to read it. New paper with Rakshit Trivedi and Dylan Hadfield-Menell.
Building AI for the Democratic Matrix: A Technical Research Agenda for Normative Competence and Normative Institutions
To maintain democratic resilience, it is essential to build AI agents capable of choosing behaviors that mirror those of the human agents that constitute human democracies.
knightcolumbia.org
Rebuttal season is here, yey🤞 With many asking me, I compiled the most common misconceptions Hope the tips help 🧵 All tips: docs.google.com/document/d/14Wax8M5w8F_8miDlYJ9-I6wqpelxlXjCEUbkNzNMqqE/edit?tab=t.0#heading=h.rfq27f356vmm #AI 🤖📈🧠
New important (I hope) resource for academics working in this area.
🚀Introducing 𝐆𝐔𝐈𝐃𝐄-𝐋𝐋𝐌: A reporting checklist for using LLMs in behavioral & social science ✅GUIDE-LLM is a reporting checklist designed by 80+ experts to improve transparency, reproducibility & ethical accountability of LLM-based research 📄 llm-checklist.com
For a recent lab meeting, I wrote up a grab bag of ways to think about your development as a researcher during a PhD: emerge-lab.github.io/papers/an-un... Sharing in case folks find it useful or have feedback!
emerge-lab.github.io
🥁🥁🥁 Newly out from us today in Science Advances: “Biased AI Writing Assistants Shift Users’ Attitudes on Societal Issues”. Large Language Models are providing users with autocomplete writing suggestions on many platforms. Could these suggestions shift users’ own attitudes? (spoiler: YES) (1/7)
There's been a lot of excitement about pluralistic value alignment 🌈 — AI that reflects the full range of human perspectives But no formal way to benchmark whether we're actually making progress. 🤔 Introducing 𝐎𝐕𝐄𝐑𝐓𝐎𝐍𝐁𝐄𝐍𝐂𝐇. 🎉Accepted to #ICLR2026 1/n 🧵
Do LLMs Benefit from Their Own Words?🤔 In multi-turn chats, models are typically given their own past responses as context. But do their own words always help… Or are they more often a waste of compute and a distraction? 🧵 arxiv.org/abs/2602.24287
Models are now expert math solvers, and so AI for math education is receiving increasing attention. Our new preprint evaluates 11 VLMs on our QA benchmark, DrawEduMath. We highlight a startling gap: models perform less well on inputs from K-12 students who need more help. 🧵
Yesterday was my last day at MSR. We recently learned that our roles were eliminated, and with them our little FATE Montreal team. I joined MSR a bit over 7.5 years ago while on active chemotherapy, and being at MSR has overlapped with so much change in my life.