Who's attending @neuripsconf.bsky.social this year and which location? 🤖☃️ #NeurIPS2026
Hanna Wallach
@hannawallach.bsky.social
VP and Distinguished Scientist at Microsoft Research NYC. AI evaluation and measurement, responsible AI, computational social science, machine learning. She/her. One photo a day since January 2018: https://www.instagram.com/logisticaggression/
Microsoft Research NYC is hiring a researcher in the space of AI and society!
[NeurIPS '25] Our oral slot and poster session on "Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research" are tomorrow, December 4! [https://arxiv.org/abs/2412.06966] Oral: 3:30-4pm PST, Upper Level Ballroom 20AB Poster 1307: 4:30:-7:30pm PST, Exhibit Hall C-E
I'm super excited about the 20th @wimlworkshop.bsky.social, which is taking place tomorrow in San Diego, co-located with @neuripsconf.bsky.social!!! 🎉 To celebrate, @jennwv.bsky.social and I recorded a podcast episode! Check it out here: www.microsoft.com/en-us/resear...
Ideas: Community building, machine learning, and the future of AI
As the Women in Machine Learning Workshop (WiML) marks its 20th annual gathering, cofounders, friends, and collaborators Jenn Wortman Vaughan and Hanna Wallach reflect on WiML’s evolution, navigating ...
microsoft.com
On my way to San Diego for #NeurIPS and the 20th anniversary of #WiML! To celebrate, @hannawallach.bsky.social and I recorded a podcast discussing WiML’s journey, our friendship and collaborations, research we're excited about, & advice we'd give our younger selves! www.microsoft.com/en-us/resear...
Ideas: Community building, machine learning, and the future of AI
As the Women in Machine Learning Workshop (WiML) marks its 20th annual gathering, cofounders, friends, and collaborators Jenn Wortman Vaughan and Hanna Wallach reflect on WiML’s evolution, navigating ...
microsoft.com
Alright, it's that time of year: Who all is going to @neuripsconf.bsky.social this year??? #NeurIPS2025 🤖☃️
Three exciting opportunities at @msftresearch.bsky.social in NYC!!! 🎉 Internship w/ FATE: apply.careers.microsoft.com/careers/job?... Internship w/ STAC on AI evaluation and measurement: apply.careers.microsoft.com/careers/job?... Postdoc w/ FATE: apply.careers.microsoft.com/careers/job?...
FATE internships: apply.careers.microsoft.com/careers/job?... FATE postdocs: apply.careers.microsoft.com/careers/job?... And internships with our close collaborators at STAC: apply.careers.microsoft.com/careers/job?...
This is happening now!!!
If you're at @icmlconf.bsky.social this week, come check out our poster on "Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge" presented by the amazing @afedercooper.bsky.social from 11:30am--1:30pm PDT on Weds!!! icml.cc/virtual/2025...
1) (Tomorrow!) Wed 7/16, 11am-1:30 pm PT poster for "Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge" (E. Exhibition Hall A-B, E-503) Work led by @hannawallach.bsky.social + @azjacobs.bsky.social arxiv.org/abs/2502.00561
Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge
The measurement tasks involved in evaluating generative AI (GenAI) systems lack sufficient scientific rigor, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges com...
arxiv.org
If you're at @icmlconf.bsky.social this week, come check out our poster on "Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge" presented by the amazing @afedercooper.bsky.social from 11:30am--1:30pm PDT on Weds!!! icml.cc/virtual/2025...
ICML Poster Position: Evaluating Generative AI Systems Is a Social Science Measurement ChallengeICML 2025
icml.cc
Generative language systems are everywhere, and many of them stereotype, demean, or erase particular social groups.
Alright, people, let's be honest: GenAI systems are everywhere, and figuring out whether they're any good is a total mess. Should we use them? Where? How? Do they need a total overhaul? (1/6)
I'm so excited this paper is finally online!!! 🎉 We had so much fun working on this with @emmharv.bsky.social!!! Thread below summarizing our contributions...
📣 "Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems" is forthcoming at #ACL2025NLP - and you can read it now on arXiv! 🔗: arxiv.org/pdf/2506.04482 🧵: ⬇️
Exciting news: The Fairness, Accountability, Transparency and Ethics (FATE) group at Microsoft Research NYC is hiring a predoctoral fellow!!! 🎉 www.microsoft.com/en-us/resear...
FATE Research Assistant (“Pre-doc”) - Microsoft Research
The Fairness, Accountability, Transparency, and Ethics (FATE) Research group at Microsoft Research New York City (MSR NYC) is looking for a pre-doctoral research assistant (pre-doc) to start August 20...
microsoft.com
Exciting news!!! This just got into @icmlconf.bsky.social as a position paper!!! 🎉 More updates to come as we work on the camera-ready version!!!
Remember this @neuripsconf.bsky.social workshop paper? We spent the past month writing a newer, better, longer version!!! You can find it online here: arxiv.org/abs/2502.00561
Reading - Evaluating Evaluations for GenAI from @hannawallach.bsky.social madesai.bsky.social afedercooper.bsky.social et al-This work dovetails with our work at @worldprivacyforum.bsky.social on measuring AI governance tools from governments, through privacy/ policy lens arxiv.org/pdf/2502.00561
At the #HEAL workshop, I'll present "Systematizing During Measurement Enables Broader Stakeholder Participation" on the ways we can further structure LLM evaluations and open them for deliberation. A project led by @hannawallach.bsky.social
2. Also Saturday, @amabalayn.bsky.social will represent our piece arguing that systematization during measurement enables broad stakeholder participation in AI evaluation. This came out of a huge group collaboration led by @hannawallach.bsky.social: bsky.app/profile/hann... heal-workshop.github.io
📣 New paper! The field of AI research is increasingly realising that benchmarks are very limited in what they can tell us about AI system performance and safety. We argue and lay out a roadmap toward a *science of AI evaluation*: arxiv.org/abs/2503.05336 🧵
This link will take you to a page that’s not on LinkedIn
lnkd.in
⚫⚪ It's coming...SHADES. ⚪⚫ The first ever resource of multilingual, multicultural, and multigeographical stereotypes, built to support nuanced LLM evaluation and bias mitigation. We have been working on this around the world for almost **4 years** and I am thrilled to share it with you all soon.
Remember this @neuripsconf.bsky.social workshop paper? We spent the past month writing a newer, better, longer version!!! You can find it online here: arxiv.org/abs/2502.00561
Position: Evaluating Generative AI Systems is a Social Science Measurement Challenge
The measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons...
arxiv.org
Evaluating Generative AI Systems is a Social Science Measurement Challenge: arxiv.org/abs/2411.10939 TL;DR: The ML community would benefit from learning from and drawing on the social sciences when evaluating GenAI systems.
🚨Postdoc Alert 🚨 The Computational Social Science group at Microsoft Research NYC (Jake Hofman, David Rothschild, Dan Goldstein) is hiring a postdoc! jobs.careers.microsoft.com/global/en/jo... Deadline: December 20, 2024
Search Jobs | Microsoft Careers
jobs.careers.microsoft.com
Microsoft's Computational Social Science group may have the opportunity to hire one researcher Senior: 0-3 yrs post PhD jobs.careers.microsoft.com/global/en/jo... Principal: 3+ yrs post PhD jobs.careers.microsoft.com/global/en/sh... Please note: our ability to hire this season is not certain
It's company holiday party season! Every year I start a thread of my favorite questions guaranteed to get you 20 minutes of lively conversation (as an introvert, this is how I thrive at parties). What are your favorites? Here are some of mine...
"there's a lot of qualitative work that goes into designing quantitative metrics" -- @azjacobs.bsky.social "how do we translate between benchmark performance and what it will really be like to use a model" -- Su Lin Blodgett
Super interesting panel discussion taking place right now at the Evaluating Evaluations workshop at @neuripsconf.bsky.social with amazing panelists @abeba.bsky.social, @azjacobs.bsky.social, Su Lin Blodgett, and Lee Wan Sie!!! #NeurIPS2024
Super interesting panel discussion taking place right now at the Evaluating Evaluations workshop at @neuripsconf.bsky.social with amazing panelists @abeba.bsky.social, @azjacobs.bsky.social, Su Lin Blodgett, and Lee Wan Sie!!! #NeurIPS2024
And, as if that wasn't enough excitement for one day, we'll also be presenting a poster on "Red Teaming: Everything Everywhere All at Once" at the Safe GenAI @neuripsconf.bsky.social workshop from 3--5pm: neurips.cc/virtual/2024... #NeurIPS2024
Safe Generative AINeurIPS 2024
neurips.cc
Super excited for the Evaluating Evaluations workshop at @neuripsconf.bsky.social today!!! evaleval.github.io #NeurIPS2024 @msftresearch.bsky.social's FATE group, Sociotechnical Alignment Center, and friends will be presenting several papers there. See below for details...
Home - EvalEval 2024
A NeurIPS 2024 workshop on best practices for measuring the broader impacts of generative AI systems
evaleval.github.io
Stop by the 345--430pm poster session at the Statistical Frontiers in LLMs @neuripsconf.bsky.social workshop today (Dec 14) to catch posters on the following papers... #NeurIPS2024
New paper on why machine "unlearning" is much harder than it seems is now up on arXiv: arxiv.org/abs/2412.06966 This was a huuuuuge cross-disciplinary effort led by @msftresearch.bsky.social FATE postdoc @grumpy-frog.bsky.social!!!
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy, Research, and Practice
We articulate fundamental mismatches between technical methods for machine unlearning in Generative AI, and documented aspirations for broader impact that these methods could have for law and policy. ...
arxiv.org