Public datasets of Chest X-Rays (CXRs) are a popular benchmark for ML computer vision models, which often report strong average-case performance. But does this aggregate performance reflect actual utility in the clinic?
Michael Oberst
@moberst.bsky.social
Assistant Prof. of CS at Johns Hopkins Visiting Scientist at Abridge AI Causality & Machine Learning in Healthcare Prev: PhD at MIT, Postdoc at CMU
Today I am very proud to announce the release of the Antibiotic Resistance Microbiology Dataset - Mass General Brigham (ARMD-MGB; physionet.org/content/armd...) , as part of an NIH-funded collaboration led by Jonathan Chen at Stanford. (1/6)
Antibiotic Resistance Microbiology Dataset Mass General Brigham (ARMD-MGB) v1.0.0
ARMD-MGB contains detailed microbiology and clinical metadata for >225,000 patients and >970,000 cultures collected over 10 years
physionet.org
Come join my group at Johns Hopkins! I'm recruiting CS PhD students for Fall'26 (deadline: Dec 15) who are interested in safe/reliable AI in healthcare. See my website (link in reply) for more info. I'm also headed to #NeurIPS, and happy to chat with prospective students!
Randomized trials (RCTs) help evaluate if deploying AI/ML systems actually improves outcomes (e.g., survival rates in a healthcare context). But AI/ML systems can change: Do we need a new RCT every time we update the model? Not necessarily, as we show in our UAI paper! arxiv.org/abs/2502.09467
In this conversation I have been endorsed as "twee" and "not a crank". BTW, I'm on the job market this year. If you are interested hiring an economist in macro/metrics/computational/ML with such stellar endorsements, please get in touch!
Oh I have no strong view on the substance, I just missed conversations that go like this (not with cranks) and I'm happy there's a platform where they're happening again.
I'm recruiting PhD students for Fall 2025! CS PhD Deadline: Dec. 15th. I work on safe/reliable ML and causal inference, motivated by healthcare applications. Beyond myself, Johns Hopkins has a rich community of folks doing similar work. Come join us!
Medically adapted foundation models (think Med-*) turn out to be more hot air than hot stuff. Correcting for fatal flaws in evaluation, the current crop are no better on balance than generic foundation models, even on the very tasks for which benefits are claimed. arxiv.org/abs/2411.04118
Medical Adaptation of Large Language and Vision-Language Models: Are We Making Progress?
Several recent works seek to develop foundation models specifically for medical applications, adapting general-purpose large language models (LLMs) and vision-language models (VLMs) via continued pret...
arxiv.org
I am recruiting PhD students at Duke! Please apply to Duke CS or CBB if you are interested in developing new methods and paradigms for NLP/LLMs in healthcare. For details, see here: monicaagrawal.com/home/researc....
Couldn't find a machine learning for health starter pack so I made one. DM/Reply if you want to be added! go.bsky.app/PJKJ8vK
First post! I'm recruiting PhD students this PhD admission cycle who want to work on: a) impactful ML methods for healthcare 🤖, b) computational methods to improve health equity ⚖️, or c) AI for women's health or climate health 🤰🌎 Apply via UC Berkeley CPH or EECS (AI-H) 🌉. irenechen.net/join-lab/