Joachim Baumann

@joachimbaumann.bsky.social

Postdoc @stanfordnlp.bsky.social / previously @milanlp.bsky.social / Computational social science, LLMs, algorithmic fairness

Our work got featured in Science News! What a great end to an amazing #ICML2026🇰🇷 Peer review is at a crossroads. How we handle the submission and review crisis over the next years will shape the future of scientific publishing. True impact will come from people writing fewer, more meaningful papers!

Science News@sciencenews.bsky.social · 4w ago

AI tools might inadvertently perpetuate the biases they’re known to carry and reduce the variety of opinions weighing in on new science. https://www.sciencenews.org/article/ai-tools-science-peer-review-problems

Just arrived at ICML 🇰🇷😍 Get up early tomorrow to hear me talk about how (not) to solve the peer review crisis, or find me at one of my poster presentations. Paper links: ✅ AI Peer Review: arxiv.org/abs/2605.03202 ✅ SWE-chat: arxiv.org/pdf/2604.20779

ICML Conference schedule with two papers. Left, "Stop Automating Peer Review Without Rigorous Evaluation": Oral presentation, Wed 7/8/2026, 10:00–10:15 AM KST, Grand Ballroom 101–105; Poster, Wed 7/8/2026, 2:30–4:15 PM KST, Hall A #3003. Right, "SWE-chat": at the 5th Deep Learning for Code Workshop, Fri 7/10/2026, 13:00–14:30 KST, Hall B2.
Joachim Baumann@joachimbaumann.bsky.social · 3mo ago

Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026

First page of the ICML 2026 spotlight paper "Stop Automating Peer Review Without Rigorous Evaluation" by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University). The abstract argues that today's AI systems should not be used to produce paper reviews, grounded in two empirical findings: a "hivemind effect" where AI reviewers show excessive agreement and reduce perspective diversity, and "paper laundering," where prompting an LLM to rewrite a paper trivially increases AI reviewer scores through stylistic changes rather than scientific improvements. The paper calls for a science of peer review automation rather than wholesale deployment of general-purpose LLMs.

The dataset we call GoogleTrendArchive has over 7 million trend episodes spanning over 1 year since Nov. 28th 2024. We cover all 1358 locations available. Direct link to the dataset huggingface.co/datasets/aur... Joint work with Anikó Hannák @scg-uzh.bsky.social and @joachimbaumann.bsky.social

aurman/GoogleTrendArchive · Datasets at Hugging Face

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

huggingface.co

Working with web search data? Ever wanted to get access to historical data on what was trending in different locations - beyond the 7 days of such history that Google's Trending Now provides? We've got you covered with the new dataset paper, accepted at ICWSM, preprint here arxiv.org/abs/2603.21871

GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide

GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 locations. Unlike Google Tre...

arxiv.org

Did you know that from tomorrow, Qualtrics is offering synthetic panels (AI-generated participants)? Follow me down a rabbit hole I'm calling "doing science is tough and I'm so busy, can't we just make up participants?"

Text reads: About synthetic panels
Recruiting the right participants for a study can be difficult. You may not get the exact demographics you need, and the shorter the deadline, the less sure you can be that everyone will answer on time. One possible solution can be to use synthetic panels.

Synthetic panels are powered by a first party proprietary AI model developed here at Qualtrics. Our synthetic panel is trained on thousands of responses from a variety of demographic backgrounds in order to more accurately predict how certain populations would respond to a survey.

Our synthetic panel is based on the United States General Population, and is only available in English. This panel comes with ready-made quotas and target breakouts in order to represent your chosen population and make it easy to launch your survey right away.Text reads:
Question-writing best practices
To get the most reliable and actionable results from synthetic audiences, consider these question-writing best practices:

Ask forward-looking and attitudinal questions.
Synthetic panels perform best with perceptions, preferences, and intent-based questions. For example, “How likely are you to try…?”
Synthetic panels are less applicable for studies on past behaviors, detailed recall, brand recall, or awareness questions. For example, “When did you last visit…?”Text reads:
Discussion
The current study aimed to conduct a meta-analysis of the TPB when applied to health behaviours which addressed the limitations of previous reviews by including only prospective tests of behaviour, applying RE meta-analytic procedures, correcting correlations for sampling and measurement error, and hierarchically analysing the effect of behaviour type and sample and methodological moderators. Some 237 tests were identified which examined relations amongst model components. Overall the analysis indicated that the TPB could explain 19.3% of the variance in behaviour and 44.3% of the variance in intention across studies. This level of prediction of behaviour is slightly lower than that of previous meta-analytic reviews which have found between 27% (Armitage & Conner, 2001; Hagger et al., 2002) and 36% (Trafimow et al., 2002)
of the variance in behaviour to be explained by intention and PBC.

Google AI overviews now reach over 2B users worldwide. But how reliable are they on high stakes topics - for instance, pregnancy and baby care? We have a new paper - led by Desheng Hu, now accepted at @icwsm.bsky.social - exploring that and finding many issues Preprint: arxiv.org/abs/2511.12920 🧵👇

Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy

Google Search increasingly surfaces AI-generated content through features like AI Overviews (AIO) and Featured Snippets (FS), which users frequently rely on despite having no control over their presen...

arxiv.org

Can AI simulate human behavior? 🧠 The promise is revolutionary for science & policy. But there’s a huge "IF": Do these simulations actually reflect reality? To find out, we introduce SimBench: The first large-scale benchmark for group-level social simulation. (1/9)

If you feel uneasy using LLMs for data annotation, you are right (if not, you should). It offers new chances for research that is difficult with traditional #NLP/#textasdata methods, but the risk of false conclusions is high! Experiment + *evidence-based* mitigation strategies in this preprint 👇

Joachim Baumann@joachimbaumann.bsky.social · 11mo ago

🚨 New paper alert 🚨 Using LLMs as data annotators, you can produce any scientific result you want. We call this **LLM Hacking**. Paper: arxiv.org/pdf/2509.08825

We present our new preprint titled "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation".
We quantify LLM hacking risk through systematic replication of 37 diverse computational social science annotation tasks.
For these tasks, we use a combined set of 2,361 realistic hypotheses that researchers might test using these annotations.
Then, we collect 13 million LLM annotations across plausible LLM configurations.
These annotations feed into 1.4 million regressions testing the hypotheses. 
For a hypothesis with no true effect (ground truth $p > 0.05$), different LLM configurations yield conflicting conclusions.
Checkmarks indicate correct statistical conclusions matching ground truth; crosses indicate LLM hacking -- incorrect conclusions due to annotation errors.
Across all experiments, LLM hacking occurs in 31-50\% of cases even with highly capable models.
Since minor configuration changes can flip scientific conclusions, from correct to incorrect, LLM hacking can be exploited to present anything as statistically significant.