Stefan Feuerriegel

@sfeuerriegel.bsky.social

Professor at LMU Munich | Institute of AI in Management ++ We develop artificial intelligence for impact ++ https://www.som.lmu.de/ai/en/index.html

โ˜€๏ธNew preprint: Hyperlocal monitoring of urban activity responses to heat We use >100M anonymized activity observations across 10 German cities to show how heat changes where people spend time ๐Ÿ‘‰informs targeted urban heat adaptation arxiv.org/abs/2607.07417

Hyperlocal monitoring of urban activity reveals responses to heat exposure

Rising temperatures create new challenges for local heat adaptation. Yet, it remains unclear how urban activity changes during hot periods and which urban environments people concentrate in as tempera...

arxiv.org

[1/4] ๐Ÿ‡ฐ๐Ÿ‡ท Excited to present our new paper โ€œFrequentist Consistency of Prior-Data Fitted Networks for Causal Inferenceโ€ at ICML 2026 in Seoul. The paper asks a simple question: when PFNs are used for causal inference, do their uncertainty estimates actually agree with classical frequentist inference?

Bild

๐ŸŽ‰ New preprint: OncoSynth ๐Ÿ‘‰We introduce a causally-aware generative ML framework for synthetic oncology cohorts โœ… Preserves structure (covariateโ†’treatmentโ†’survival) to improve treatment effect estimation Preprint: arxiv.org/abs/2606.25762

OncoSynth: Synthetic data generation for treatment effect estimation in oncology

In oncology, access to patient-level data is often restricted. Synthetic data provides an alternative for analyzing treatment effectiveness, but existing methods for synthetic data generation fail to ...

arxiv.org

Some new work on AI & science: A reporting checklist for LLMs in behavioral science Thanks to @sfeuerriegel.bsky.social for spearheading. As LLMs become more common in research, transparency is essential. We developed a reporting checklist for reporting how LLMs were used - please share! 1/4

A reporting checklist for large language models in behavioural science - Nature Human Behaviour

Large language models offer new opportunities for behavioural science, but their rapid evolution poses challenges for research rigour. We introduce a consensus-based reporting checklist to improve tra...

nature.com

๐ŸŽ‰ New preprint: Can #LLMs automate reproducibility checks? ๐Ÿ‘‰We test this on 76 social & behavioral studies โœ…Our LLM pipeline arrived at original qualitative conclusions in 96% of cases ๐ŸŽฏWe recovered the effect sizes in 41% (human reanalysts: 34%) arxiv.org/abs/2606.13670

Automated reproducibility assessments in the social and behavioral sciences using large language models

Reproducibility in the social and behavioral sciences is typically evaluated by independent researchers who reanalyze the original data to assess whether the published findings can be recovered. Howev...

arxiv.org

โšฝ๏ธIntroducing ๐‹๐‹๐Œ ๐’๐จ๐œ๐œ๐ž๐ซ๐€๐ซ๐ž๐ง๐š: a benchmarking platform for real-world LLM predictions โœ…Predictions for 104 games @FIFAWorldCup with frontier LLMs ๐Ÿ†Opus 4.8: Spain / GPT-5.5: Spain / Grok 4.3: Argentina / Mistral-large: France โžก๏ธAll predictions & leaderboard llmsoccerarena.up.railway.app

LLM SoccerArena

Compare football score predictions from multiple LLMs.

llmsoccerarena.up.railway.app

๐Ÿง  #AI can improve diagnostic accuracy in radiology โ€“ but only if its reasoning is transparent. A new #LMU study found that step-by-step explanations helped physicians better assess and critically evaluate AI recommendations, leading to significantly higher accuracy. ๐Ÿ“ˆ #LMUMunich #Radiology

Medical diagnoses: how AI explanations help doctors

A team led by LMU researcher Stefan Feuerriegel shows that AI models like ChatGPT can improve diagnostic accuracy in radiology โ€“ but how the AI explains its recommendations is crucial.

lmu.de

๐Ÿ“ฃNew: ๐€๐๐š๐ฉ๐ญ๐ข๐ฏ๐ž ๐„๐ฑ๐ฉ๐ž๐ซ๐ข๐ฆ๐ž๐ง๐ญ๐š๐ญ๐ข๐จ๐ง ๐Ÿ๐จ๐ซ ๐‚๐ž๐ง๐ฌ๐จ๐ซ๐ž๐ ๐’๐ฎ๐ซ๐ฏ๐ข๐ฏ๐š๐ฅ ๐Ž๐ฎ๐ญ๐œ๐จ๐ฆ๐ž๐ฌ ๐Ÿ‘‰We develop ASE: an adaptive experimentation framework for survival outcomes with censoring โ˜‘๏ธachieves A-optimal semiparametric efficiency bound ๐Ÿ“„ arxiv.org/abs/2605.18459

Adaptive Experimentation for Censored Survival Outcomes

Adaptive experimentation enables efficient estimation of causal effects, but existing methods are not designed for survival data with censoring, where event times are only partially observed (e.g., ov...

arxiv.org

๐ŸŽ‰ New paper: ๐’๐ค๐ข๐ฅ๐ฅ๐†๐ž๐ง โ€” Verified inference-time agent skill synthesis ๐Ÿ‘‰ Skills can improve LLM agents โ€” but are still mostly hand-written. We introduce #SkillGen: a multi-agent framework that learns skills from trajectories using contrastive induction + a generateโ€“verifyโ€“refine loop. ๐Ÿงต1/3

Bild

๐Ÿš€New paper @ npj digital medicine: ๐“๐ก๐ž ๐ž๐Ÿ๐Ÿ๐ž๐œ๐ญ ๐จ๐Ÿ ๐ฆ๐ž๐๐ข๐œ๐š๐ฅ ๐ž๐ฑ๐ฉ๐ฅ๐š๐ง๐š๐ญ๐ข๐จ๐ง๐ฌ ๐Ÿ๐ซ๐จ๐ฆ ๐ฅ๐š๐ซ๐ ๐ž ๐ฅ๐š๐ง๐ ๐ฎ๐š๐ ๐ž ๐ฆ๐จ๐๐ž๐ฅ๐ฌ ๐จ๐ง ๐๐ข๐š๐ ๐ง๐จ๐ฌ๐ญ๐ข๐œ ๐๐ž๐œ๐ข๐ฌ๐ข๐จ๐ง๐ฌ ๐ข๐ง ๐ซ๐š๐๐ข๐จ๐ฅ๐จ๐ ๐ฒ ๐Ÿ”What is the right format to present explanations in medical #LLM? ๐Ÿ‘‰We ran a large experiment with 101๐Ÿ‘ฉโ€โš•๏ธ(N=2020 diagnostic decisions) www.nature.com/articles/s41...

The effect of medical explanations from large language models on diagnostic accuracy in radiology - npj Digital Medicine

npj Digital Medicine - The effect of medical explanations from large language models on diagnostic accuracy in radiology

nature.com