Michelle L. Ding

@michelleding.bsky.social

researcher/organizer critically investigating how AI systems impact communities. cs phd @ brown cntr. she/her. 🌷 https://michelle-ding.github.io/ 💭 https://michellelding.substack.com/

Most Multilingual benchmarks measure what models know, not what they can reliably do. Our #ACL2026 paper introduces functional benchmarks in six languages from English to Yoruba to test whether models can actually execute tasks across languages, not just answer fixed questions about them. 🧵1/n

Bild

Anthropic just shut access to Fable & Mythos, their most powerful models. I wrote up what happened + the policy and psychodrama behind it (with all the links & receipts) here — the thread below is the TL;DR 🧵 blog.geomblog.org/2026/06/the-...

The "fable" of Anthropic and the USG

News moves fast. 12 hours ago I was enjoying the demolition that the US put on Paraguay when I heard that Anthropic had shut down access to ...

blog.geomblog.org

Tech Policy Press fellow Petra Molnar highlights the AI Resist List: a global database documenting acts of resistance to the AI industry. From legal challenges and worker organizing to artistic interventions, the project seeks to challenge the “scale at all costs” development of AI.

The World Is Already Resisting AI. Now, There is a List to Prove It.

Petra Molnar spotlights the launch of the AI Resist List, documenting global challenges to AI expansion.

buff.ly

Very glad to be a part of a new paper detailing how developers and developer platforms can prevent AIG-NCII, a form of image based sexual abuse that disproportionately harms women and girls. Thanks to all the collaborators and Max Kamachee & @scasper.bsky.social for leading this important project!

Cas (Stephen Casper)@scasper.bsky.social · 8mo ago

Did you know that one base model is responsible for 94% of model-tagged NSFW AI videos on CivitAI? This new paper studies how a small number of models power the non-consensual AI video deepfake ecosystem and why their developers could have predicted and mitigated this.

ACM members/computing researchers who should be members interested in contributing should join the subcommittee's mailing list! One of our goals here is to build policy coalitions across institutions so we can do more as a collective 💪 and balance special interest groups.

Suresh Venkatasubramanian @geomblog.bsky.social · 8mo ago

@reniebird.bsky.social and I have just been appointed to co-Chair @TheOfficialACM's US Technology Policy Committee’s Subcommittee on AI and Algorithms. cs.brown.edu/news/2025/11...

PSA: tips to protect yourself from scams on Signal. Every major comms platform has to contend w phishing, impersonation, & scams. Sadly. Signal is major, and as we've grown we've heard about more of these attacks--scammy people pretending to be something or someone to trick and abuse others. 1/

I wrote a (personal) blog post about my hopes and dreams for AI policy, my devastation after the US Election, and my process of picking myself off the floor by rebuilding an optimistic vision for AI scientists in government through education: simons.berkeley.edu/news/rebuild...

Rebuilding an Optimistic Vision for AI Policy

Recall November 6, 2024 — the day after the U.S. election. I was driving back to my home in Washington, DC, from Ohio with colleagues. I was heartbroken not because of the rebuke to my political party...

simons.berkeley.edu

Have you or a loved one been misgendered by an LLM? How can we evaluate LLMs for misgendering? Do different evaluation methods give consistent results? Check out our preprint led by the newly minted Dr. @arjunsubgraph.bsky.social, and with Preethi Seshadri, Dietrich Klakow, Kai-Wei Chang, Yizhou Sun

Agree to Disagree? A Meta-Evaluation of LLM Misgendering

Numerous methods have been proposed to measure LLM misgendering, including probability-based evaluations (e.g., automatically with templatic sentences) and generation-based evaluations (e.g., with aut...

arxiv.org

Hi #COLM2025! 🇨🇦 I will be presenting a talk on the importance of community-driven LLM evaluations based on an opinion abstract I wrote with Jo Kavishe, @victorojewale.bsky.social and @geomblog.bsky.social tomorrow at 9:30am in 524b for solar-colm.github.io Hope to see you there!

Third Workshop on Socially Responsible Language Modelling Research (SoLaR) 2025

COLM 2025 in-person Workshop, October 10th at the Palais des Congrès in Montreal, Canada

solar-colm.github.io

I'll be presenting a position paper about consumer protection and AI in the US at ICML. I have a surprisingly optimistic take: our legal structures are stronger than I anticipated when I went to work on this issue in Congress. Is everything broken rn? Yes. Will it stay broken? That's on us.

A poster for the paper "Position: Strong Consumer Protection is an Inalienable Defense for AI Safety in the United States"

With their 'Sovereignty as a Service' offerings, tech companies are encouraging the illusion of a race for sovereign control of AI while being the true powers behind the scenes, write Rui-Jie Yew, Kate Elizabeth Creasey, Suresh Venkatasubramanian.

'Sovereignty' Myth-Making in the AI Race | TechPolicy.Press

Tech companies stand to gain by encouraging the illusion of a race for 'sovereign' AI, write Rui-Jie Yew, Kate Elizabeth Creasey, Suresh Venkatasubramanian.

techpolicy.press