Dang Nguyen

@divingwithorcas.bsky.social

Computer Science PhD student at UChicago | Member of the Chicago Human+AI lab @chicagohai.bsky.social

AI can accelerate scientific discovery, but only if we get the scientist–AI interaction right. The dream of “autonomous AI scientists” is tempting: machines that generate hypotheses, run experiments, and write papers. But science isn’t just automation. cichicago.substack.com/p/the-mirage... 🧵

The Mirage of Autonomous AI Scientists

Science as AI’s killer application cannot succeed without scientist-AI interaction: Introducing Hypogenic.ai.

cichicago.substack.com

📣 Announcing our poster session at COLM 2025: On the Effectiveness and Generalization of Race Representations for Debiasing High-Stakes Decisions I will talk about biases in LLMs and how to mitigate them. Come say hi! Poster #43, 4:30 PM

Bild
Dang Nguyen@divingwithorcas.bsky.social · last yr.

1/n You may know that large language models (LLMs) can be biased in their decision-making, but ever wondered how those biases are encoded internally and whether we can surgically remove them?

HR Simulator™: a game where you gaslight, deflect, and “let’s circle back” your way to victory. Every email a boss fight, every “per my last message” a critical hit… or maybe you just overplayed your hand 🫠 Can you earn Enlightened Bureaucrat status? (link below!)

Bild

Prompting is our most successful tool for exploring LLMs, but the term evokes eye-rolls and grimaces from scientists. Why? Because prompting as scientific inquiry has become conflated with prompt engineering. This is holding us back. 🧵and new paper with @ari-holtzman.bsky.social .

Bild

🚨 New paper alert 🚨 Ever asked an LLM-as-Marilyn Monroe who the US president was in 2000? 🤔 Should the LLM answer at all? We call these clashes Concept Incongruence. Read on! ⬇️ 1/n 🧵

Bild

1/n 🚀🚀🚀 Thrilled to share our latest work🔥: HypoEval - Hypothesis-Guided Evaluation for Natural Language Generation! 🧠💬📊 There’s a lot of excitement around using LLMs for automated evaluation, but many methods fall short on alignment or explainability — let’s dive in! 🌊

🧑‍⚖️How well can LLMs summarize complex legal documents? And can we use LLMs to evaluate? Excited to be in Albuquerque presenting our paper this afternoon at @naaclmeeting 2025!

Bild

🚀🚀🚀Excited to share our latest work: HypoBench, a systematic benchmark for evaluating LLM-based hypothesis generation methods! There is much excitement about leveraging LLMs for scientific hypothesis generation, but principled evaluations are missing - let’s dive into HypoBench together.

Bild

Encourage your students to submit posters and register! Limited free housing is provided for student participants only, on a first-come (i.e., request)-first-serve basis. We are also actively looking for sponsors. Reach out if you are interested! Please repost! Help spread the words!

@chenhaotan.bsky.social · last yr.

The Midwest Machine Learning Symposium will happen in Chicago on June 23-4 on the University of Chicago campus (midwest-ml.org/2025/). We have an amazing lineup of speakers:@profsanjeevarora.bsky.social from Princeton, Heng Ji from UIUC, Tuomas Sandholm from CMU, @ravenben.bsky.social from UChicago.