xavier roberts-gaal

@xrg.bsky.social

three language models in a trench coat harvard psych (scholar.harvard.edu/xrg)

I have a new paper. We look at ~all stats articles in political science post-2010 & show that 94% have abstracts that claim to reject a null. Only 2% present only null results. This is hard to explain unless the research process has a filter that only lets rejections through.

It must be very hard to publish null results
Publication practices in the social sciences act as a filter that favors statistically significant results over null findings. While the problem of selection on significance (SoS) is well-known in theory, it has been difficult to measure its scope empirically, and it has been challenging to determine how selection varies across contexts. In this article, we use large language models to extract granular and validated data on about 100,000 articles published in over 150 political science journals from 2010 to 2024. We show that fewer than 2% of articles that rely on statistical methods report null-only findings in their abstracts, while over 90% of papers highlight significant results. To put these findings in perspective, we develop and calibrate a simple model of publication bias. Across a range of plausible assumptions, we find that statistically significant results are estimated to be one to two orders of magnitude more likely to enter the published record than null results. Leveraging metadata extracted from individual articles, we show that the pattern of strong SoS holds across subfields, journals, methods, and time periods. However, a few factors such as pre-registration and randomized experiments correlate with greater acceptance of null results. We conclude by discussing implications for the field and the potential of our new dataset for investigating other questions about political science.

🚨New Preprint: We develop a novel task that probes counterfactual thinking without using counterfactual language, and that teases apart genuine counterfactual thinking from related forms of thinking. Using this task, we find that the ability for counterfactual thinking emerges around 5 years of age.

Bild

🚨 New preprint 🚨 Across 3 experiments (n = 3,285), we found that interacting with sycophantic (or overly agreeable) AI chatbots entrenched attitudes and led to inflated self-perceptions. Yet, people preferred sycophantic chatbots and viewed them as unbiased! osf.io/preprints/ps... Thread 🧵

Abstract and results summary

love this really elegant paper spearheaded by Linas! one of the clearest instances of resource-rational social cognition i've seen worth a read!

Linas Nasvytis@linasnasvytis.bsky.social · 11mo ago

🚨New paper out w/ @gershbrain.bsky.social & @fierycushman.bsky.social from my time @Harvard! Humans are capable of sophisticated theory of mind, but when do we use it? We formalize & document a new cognitive shortcut: belief neglect — inferring others' preferences, as if their beliefs are correct🧵

💙New paper!💙 How is knowledge transmitted across generations in a foraging society? With @danielredhead.bsky.social we found: In BaYaka foragers, long-term skills pass in smaller, sparser networks, while short-term food info circulates broadly & reciprocally academic.oup.com/pnasnexus/ar...

Transmission networks of long-term and short-term knowledge in a foraging society

Abstract. Cultural transmission across generations is key to cumulative cultural evolution. While several mechanisms—such as vertical, horizontal, and obli

academic.oup.com

We often hear from reviewers: "what about demand effects?" So we developed a method to eliminate them. Something weird happened during testing: We couldn’t detect demand effects in the first place! (1/8)

Summary of design and results from our three studies. (A: Design) Each study used a similar experimental design, measuring both positive and negative demand in an online experiment, with three commonly-used task types (dictator game, vignette, intervention). Our experiments had ns ≈ 250 per cell. (B: Results) Observed demand effects were statistically indistinguishable from zero. The plot shows means and 95% confidence intervals for standardized mean differences derived from frequentist analyses of each experiment and an inverse variance-weighted fixed-effect estimator pooling all experiments (solid bars). Prior measurements of experimenter demand from a previous dictator game experiment (de Quidt et al., 2018; standardized mean difference from regression coefficient) and a meta-analysis primarily including small-sample, in-person studies (Coles et al., 2025; Hedge’s g statistic) are also shown for comparison (striped bars). The main text includes Bayesian analyses that quantify our uncertainty.

🔥Exciting news in experimental philosophy🔥 Very happy to announce that there will be soon a new journal named “Experimental Philosophy”. It will be open access, free of charge for authors and follow all Open Science principles. Editors and Editorial Board below. More information coming soon...

CALL FOR PAPERS: SPP 2024 The Society for Philosophy and Psychology (SPP) invites submissions of papers to be presented at its 50th Annual Meeting to be held June 19-June 22, 2024 at Purdue University (local organizer: Corey Maley).

Delighted to share this new preprint lead by Arthur Le Pargneux. When somebody needs to "take one for the team" (pull a late night, walk an extra mile, trudge through mud), do people think that moral responsibility falls upon whoever has the weakest bargaining position? osf.io/preprints/ps... 1/2