New piece in The Atlantic! We always hear that AI will cure cancer, and I would immediately benefit if it did. Still, I argue that racing ahead on generalist AI models creates unclear benefits for cancer that are outweighed by broader societal harms. Gift link: www.theatlantic.com/technology/2...
Emma Pierson
@emmapierson.bsky.social
Assistant professor of CS at UC Berkeley, core faculty in Computational Precision Health. Developing ML methods to study health and inequality. "On the whole, though, I take the side of amazement." https://people.eecs.berkeley.edu/~emmapierson/
Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access!
IPUMS Spatial Student Award is a tie! @gsagostini.bsky.social for "Inferring Fine-Grained Migration Patterns Across the United States." (www.nature.com/articles/s41...)
The Tech Team at ACLU is hiring! We are looking for a Data Scientist with expertise in NLP and AI ethics to work on using language tech to support ACLU's mission. Come help us tackle questions about how AI systems can be carefully applied to support the public interest. www.aclu.org/careers/appl...
Careers at ACLU
Join our team! We’re looking for committed, passionate people for open roles at the ACLU.
aclu.org
Our lab, within the Berkeley EECS department, is hiring a postdoc! More info and quick application form: forms.gle/4CcESe1TGFoo... Apply by May 1! Please reshare :)
New paper: "In Your Own Words"! We: - develop a framework to identify themes in free-text survey data - show its benefits on a new dataset of how people self-describe their race, gender, and sexual orientation - release this data for research! See @jennyshwang.bsky.social's thread below :)
New paper: "In Your Own Words"! We propose a computational framework for identifying interpretable themes from free-text survey data, and demonstrate its benefits on a new dataset of self-described race, gender, and sexual orientation. 🧵1/
We have a new piece in Nature Health led by @dmshanmugam.bsky.social, @sidhikabalachandar.bsky.social, and a wonderful team of coauthors on how to move towards a world in which race is not used in clinical algorithms!
New in Nature Health: how might we move towards a world in which race is not used in clinical algorithms? We need (1) careful comparison of race-aware and race-neutral algorithms and (2) systemic efforts to address underlying disparities.
Congratulations to @gsagostini.bsky.social, whose recent Nature Comms paper releasing a fine-grained migration dataset (www.nature.com/articles/s41...) just won a student paper award at the American Association of Geographers Annual Meeting!
Inferring fine-grained migration patterns across the United States - Nature Communications
This study releases a very high-resolution migration dataset that reveals trends that shape daily life: rising moves into high-income neighborhoods, racial gaps in upward mobility, and wildfire-driven...
nature.com
Had a great time presenting our work on building MIGRATE–a new dataset of US migration–at the @geographers.bsky.social AAG Annual Meeting today. Happy to also share that we received an AAG student paper award for this work!!! Come chat if you are at #AAG26 this week. migrate.tech.cornell.edu
Our paper, "What's in My Human Feedback", received an oral presentation at ICLR! Our method automatically+interpretably identifies preferences in human feedback data; we use this to improve personalization + safety. Reach out if you have data/use cases to apply this to! arxiv.org/pdf/2510.26202
New research is offering new insight on how Americans move — all the way to the neighborhood level. A new dataset, MIGRATE, maps annual moves with 4,600‑times more detail than standard public data, revealing patterns hidden in county‑level reporting: https://bit.ly/49XSD6w
Now out in Nature Communications - we have released a migration dataset that is - 4000x more granular than existing public data - highly correlated with Census data - being used by >100 academic, govt, and non-profit teams all over the world See @gsagostini.bsky.social's thread!
Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9
We have a new paper in JAMA Internal Medicine! Patient race is widely used in medical algorithms...but it's unclear how patients feel about this. We conduct the first nationally-representative YouGov survey to find out, producing four findings with practical clinical implications. 1/
Thanks to Kara Manke at Berkeley News for this profile of our lab's recent work on fairer decision-making in healthcare and policing! news.berkeley.edu/2026/01/20/a...
AI has a bias problem. Can we build something smarter? - Berkeley News
UC Berkeley computer scientist Emma Pierson believes we can use AI to improve our healthcare and criminal justice systems — but only if we design these algorithms with an eye toward equality.
news.berkeley.edu
We have a new paper in Science Advances proposing a simple test for bias: Is the same person treated differently when their race is perceived differently? Specifically, we study: is the same driver likelier to be searched by police when they are perceived as Hispanic rather than white? 1/
New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.
selfishly i wish we could keep divya in our lab forever but i guess it would be a disservice to the rest of the world 😅 she’s been such a wonderful mentor to me—i’ve learned a lot from how thoughtful, creative, and knowledgeable she is about everything. she’s also super funny and amazing at baking 🤭
I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.
Meeting Divya 5 years ago was one of the biggest strokes of luck in my faculty career - she is a brilliant scientist who has been foundational to so many of our lab's projects, and any institution would be lucky to hire her.
I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.
🚨 New postdoc position in our lab at Berkeley EECS! 🚨 (please reshare) We seek applicants with experience in language modeling who are excited about high-impact applications in the health and social sciences! More info in thread 1/3
📢New POSITION PAPER: Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts Despite recent results, SAEs aren't dead! They can still be useful to mech interp, and also much more broadly: across FAccT, computational social science, and ML4H. 🧵
Honored to win a #CHIL2025 best paper award for our work modeling inequality in disease progression, led by @ericachiang.bsky.social! To the NIH: health inequality remains a vital topic to support the health of all Americans. As we prove, failing to account for it biases estimates for everyone.
I can’t believe I’m saying this: our work received a Best Paper Award at #CHIL2025!! So so excited and grateful 🥰 Looking forward to day 2 of the conference with these awesome people :)
For folks at @facct.bsky.social, our very own @cornellbowers.bsky.social student @emmharv.bsky.social will present the Best-Paper-Award-winning work she led on Wednesday at 10:45 AM in the "Audit and Evaluation Approaches" session! In the meantime, 🧵 below and 🔗 here: arxiv.org/abs/2506.04419 !
A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs...
arxiv.org
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
assassinations, handcuffing a senator at press conference, marines detaining a civilian, and a military parade for the president’s birthday. rough week for democracy.
Governor Waltz has now confirmed that Hortman and her husband were killed in the attack.
and... here is the actual GIF 🙈
New work 🎉: conformal classifiers return sets of classes for each example, with a probabilistic guarantee the true class is included. But these sets can be too large to be useful. In our #CVPR2025 paper, we propose a method to make them more compact without sacrificing coverage.
The first paper of @ericachiang.bsky.social's PhD, just accepted at #CHIL2025, proposes a model of disease progression which estimates and accounts for 3 types of health disparities to more accurately measure disease severity. See her full thread below!
I’m really excited to share the first paper of my PhD, “Learning Disease Progression Models That Capture Health Disparities” (accepted at #CHIL2025)! ✨ 1/ 📄: arxiv.org/abs/2412.16406
The US government recently flagged my scientific grant in its "woke DEI database". Many people have asked me what I will do. My answer today in Nature. We will not be cowed. We will keep using AI to build a fairer, healthier world. www.nature.com/articles/d41...
My ‘woke DEI’ grant has been flagged for scrutiny. Where do I go from here?
My work in making artificial intelligence fair has been noticed by US officials intent on ending ‘class warfare propaganda’.
nature.com
A pleasure to join the Tech Policy Press podcast with @natematias.bsky.social, @geomblog.bsky.social, and @justinhendrix.bsky.social to defend the consensus that AI bias is an important concern.
Last month, a group of 200+ researchers signed a letter “Affirming the Scientific Consensus on Bias and Discrimination in AI.” It comes at a time when the Trump admin is rolling back AI policies and threatening research. Justin Hendrix spoke to three of the letter's signatories.
Lab had dogathon! Seminal dog discoveries ensued.
Our lab had a #dogathon 🐕 yesterday where we analyzed NYC Open Data on dog licenses. We learned a lot of dog facts, which I’ll share in this thread 🧵 1) Geospatial trends: Cavalier King Charles Spaniels are common in Manhattan; the opposite is true for Yorkshire Terriers.
Migration data is critical in the health, environmental, and social sciences. We're releasing a new dataset, MIGRATE: annual flows between 47 billion pairs of US Census areas. MIGRATE is: - 4600x more granular than existing public data - highly correlated with external ground-truth data 1/2
💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/
We have a new method, HypotheSAEs, for identifying *interpretable text features that predict a target variable* (aka hypothesis generation). What features of a headline predict engagement? What features of a clinical note predict whether a patient will develop cancer? 1/