Hanna Wallach

@hannawallach.bsky.social

VP and Distinguished Scientist at Microsoft Research NYC. AI evaluation and measurement, responsible AI, computational social science, machine learning. She/her. One photo a day since January 2018: https://www.instagram.com/logisticaggression/

[NeurIPS '25] Our oral slot and poster session on "Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research" are tomorrow, December 4! [https://arxiv.org/abs/2412.06966] Oral: 3:30-4pm PST, Upper Level Ballroom 20AB Poster 1307: 4:30:-7:30pm PST, Exhibit Hall C-E

⚫⚪ It's coming...SHADES. ⚪⚫ The first ever resource of multilingual, multicultural, and multigeographical stereotypes, built to support nuanced LLM evaluation and bias mitigation. We have been working on this around the world for almost **4 years** and I am thrilled to share it with you all soon.

Screenshot of 'SHADES: Towards a Multilingual Assessment of Stereotypes in Large Language Models.'
SHADES is in multiple grey colors (shades).

It's company holiday party season! Every year I start a thread of my favorite questions guaranteed to get you 20 minutes of lively conversation (as an introvert, this is how I thrive at parties). What are your favorites? Here are some of mine...

"there's a lot of qualitative work that goes into designing quantitative metrics" -- @azjacobs.bsky.social "how do we translate between benchmark performance and what it will really be like to use a model" -- Su Lin Blodgett

Hanna Wallach@hannawallach.bsky.social · 2y ago

Super interesting panel discussion taking place right now at the Evaluating Evaluations workshop at @neuripsconf.bsky.social with amazing panelists @abeba.bsky.social, @azjacobs.bsky.social, Su Lin Blodgett, and Lee Wan Sie!!! #NeurIPS2024