Max Lamparth, Ph.D.

@mlamparth.bsky.social

Research Fellow @ Stanford Intelligent Systems Laboratory and Hoover Institution at Stanford University | Focusing on interpretable, safe, and ethical AI/LLM decision-making. Ph.D. from TUM.

New job update! I’m excited to share that I’ve joined the Hoover Institution and the Stanford Intelligent Systems Laboratory (SISL) in the Stanford University School of Engineering as a Research Fellow, starting September 1st.

Bild

🚨 New paper! Medical AI benchmarks over-simplify real-world clinical practice and build on medical exam-style questions—especially in mental healthcare. We introduce MENTAT, a clinician-annotated dataset tackling real-world ambiguities in psychiatric decision-making. 🧵 Thread:

Bild

Check out our new report on multi-agent security led by Lewis Hammond and the Cooperative AI Foundation! With the deployment of increasingly agentic AI systems across domains, this research area becomes more crucial.

Bild

Submitting a benchmark to ICML? Check out our NeurIPS Spotlight paper BetterBench! We outline best practices for benchmark design, implementation & reporting to help shift community norms. Be part of the change! 🙌 + Add your benchmark to our database for visibility: betterbench.stanford.edu

Anka Reuel ➡️ NeurIPS@ankareuel.bsky.social · 2y ago

🚨 NeurIPS 2024 Spotlight Did you know we lack standards for AI benchmarks, despite their role in tracking progress, comparing models, and shaping policy? 🤯 Enter BetterBench–our framework with 46 criteria to assess benchmark quality: betterbench.stanford.edu 1/x

It was fun to contribute to this new dataset evaluating at the frontier of human expert knowledge! Beyond accuracy, the results also demonstrate the necessity for novel uncertainty quantification methods for LMs attempting challenging tasks and decision-making. Check out the paper at: lastexam.ai

Bild

In case its helpful for junior female academics, a strategy I often use when I suspect I'm getting asked to do service bc I'm female is to Suggest-A-Man. Safest to suggest someone w/roughly same seniority as you. Doesn't hurt to throw in a "They seem to have ideas on [topic of service]." 1/2

I'm seeking a postdoc to work with me and @kenholstein.bsky.social on evaluating AI/ML decision support for human experts: statmodeling.stat.columbia.edu/2024/12/10/p... P.S. I'll be at NeurIPS Thurs-Mon. Happy to talk about this position or related mutual interests! Please repost 🙏

Postdoc position at Northwestern on evaluating AI/ML decision support | Statistical Modeling, Causal Inference, and Social Science

statmodeling.stat.columbia.edu

Wait, are the AnthropicAI people seriously claiming to “unlock a rich theoretical landscape” for AI evaluation by proposing the use of…. error bars? And this secret trove of deep statistical insight starts with “use the Central Limit Theorem”? Befuddling

Bild

I'm teaching a grad seminar this winter on Prediction for Decision-making. We'll look at what it means to make good predictions for decision-making from various angles, with a focus on decisions for & about people. Reading list: statmodeling.stat.columbia.edu/2024/12/06/n... Suggestions welcome!

New Course: Prediction for (Individualized) Decision-making | Statistical Modeling, Causal Inference, and Social Science

statmodeling.stat.columbia.edu

Excited to present two papers on LM decision-making at the Harms and Risks of Military AI workshop at Mila today! It's great that the organizers created a space for the timely and crucial interdisciplinary discussions and I'm looking forward to the feedback/questions!

Bild