Emma Pierson

@emmapierson.bsky.social

Assistant professor of CS at UC Berkeley, core faculty in Computational Precision Health. Developing ML methods to study health and inequality. "On the whole, though, I take the side of amazement." https://people.eecs.berkeley.edu/~emmapierson/

Excited to see MIGRATE recognized in the IPUMS awards! Huge thanks to @emmapierson.bsky.social, @nkgarg.bsky.social, and our coauthors. Our work primarily aims to make spatiotemporal data more trustworthy and accessible to researchers, just like IPUMS. Read the paper to request data access!

IPUMS@ipums.bsky.social · 3mo ago

IPUMS Spatial Student Award is a tie! @gsagostini.bsky.social for "Inferring Fine-Grained Migration Patterns Across the United States." (www.nature.com/articles/s41...)

New paper: "In Your Own Words"! We: - develop a framework to identify themes in free-text survey data - show its benefits on a new dataset of how people self-describe their race, gender, and sexual orientation - release this data for research! See @jennyshwang.bsky.social's thread below :)

Jenny S. Wang@jennyshwang.bsky.social · 4mo ago

New paper: "In Your Own Words"! We propose a computational framework for identifying interpretable themes from free-text survey data, and demonstrate its benefits on a new dataset of self-described race, gender, and sexual orientation. 🧵1/

Now out in Nature Communications - we have released a migration dataset that is - 4000x more granular than existing public data - highly correlated with Census data - being used by >100 academic, govt, and non-profit teams all over the world See @gsagostini.bsky.social's thread!

Bild
Gabriel Agostini@gsagostini.bsky.social · 6mo ago

Our paper “Inferring fine-grained migration patterns across the United States” is now out in @natcomms.nature.com! We released a new, highly granular migration dataset. 1/9

We have a new paper in JAMA Internal Medicine! Patient race is widely used in medical algorithms...but it's unclear how patients feel about this. We conduct the first nationally-representative YouGov survey to find out, producing four findings with practical clinical implications. 1/

We have a new paper in Science Advances proposing a simple test for bias: Is the same person treated differently when their race is perceived differently? Specifically, we study: is the same driver likelier to be searched by police when they are perceived as Hispanic rather than white? 1/

Bild

New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.

Bild

selfishly i wish we could keep divya in our lab forever but i guess it would be a disservice to the rest of the world 😅 she’s been such a wonderful mentor to me—i’ve learned a lot from how thoughtful, creative, and knowledgeable she is about everything. she’s also super funny and amazing at baking 🤭

Divya Shanmugam@dmshanmugam.bsky.social · 10mo ago

I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.

Meeting Divya 5 years ago was one of the biggest strokes of luck in my faculty career - she is a brilliant scientist who has been foundational to so many of our lab's projects, and any institution would be lucky to hire her.

Divya Shanmugam@dmshanmugam.bsky.social · 10mo ago

I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.

🚨 New postdoc position in our lab at Berkeley EECS! 🚨 (please reshare) We seek applicants with experience in language modeling who are excited about high-impact applications in the health and social sciences! More info in thread 1/3

Bild

📢New POSITION PAPER: Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts Despite recent results, SAEs aren't dead! They can still be useful to mech interp, and also much more broadly: across FAccT, computational social science, and ML4H. 🧵

Bild

Honored to win a #CHIL2025 best paper award for our work modeling inequality in disease progression, led by @ericachiang.bsky.social! To the NIH: health inequality remains a vital topic to support the health of all Americans. As we prove, failing to account for it biases estimates for everyone.

Erica Chiang@ericachiang.bsky.social · last yr.

I can’t believe I’m saying this: our work received a Best Paper Award at #CHIL2025!! So so excited and grateful 🥰 Looking forward to day 2 of the conference with these awesome people :)

The first paper of @ericachiang.bsky.social's PhD, just accepted at #CHIL2025, proposes a model of disease progression which estimates and accounts for 3 types of health disparities to more accurately measure disease severity. See her full thread below!

Erica Chiang@ericachiang.bsky.social · last yr.

I’m really excited to share the first paper of my PhD, “Learning Disease Progression Models That Capture Health Disparities” (accepted at #CHIL2025)! ✨ 1/ 📄: arxiv.org/abs/2412.16406

The US government recently flagged my scientific grant in its "woke DEI database". Many people have asked me what I will do. My answer today in Nature. We will not be cowed. We will keep using AI to build a fairer, healthier world. www.nature.com/articles/d41...

My ‘woke DEI’ grant has been flagged for scrutiny. Where do I go from here?

My work in making artificial intelligence fair has been noticed by US officials intent on ending ‘class warfare propaganda’.

nature.com

Migration data is critical in the health, environmental, and social sciences. We're releasing a new dataset, MIGRATE: annual flows between 47 billion pairs of US Census areas. MIGRATE is: - 4600x more granular than existing public data - highly correlated with external ground-truth data 1/2

💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/

We have a new method, HypotheSAEs, for identifying *interpretable text features that predict a target variable* (aka hypothesis generation). What features of a headline predict engagement? What features of a clinical note predict whether a patient will develop cancer? 1/

Bild