Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...
Annika Külpmann
@anniria.bsky.social
M. Sc. psychologist, doctoral researcher (she/her) interested in meta science, psychological research methods, measurement, open science, #rstats
Psychological Science is running a special issue on "Advancing Psychological Science using Large Language Models", co-edited by me and @ruben.the100.ci. Check out the calls for submissions below! www.psychologicalscience.org/publications...
Special Issue Call for Submissions: “Advancing Psychological Science using Large Language Models”
Editors: Jamie Cummins and Ruben C. ArslanPsychological Science invites submissions for a special issue on psychological science involving large language models. We welcome theoretically and empirical...
psychologicalscience.org
RegCheck ( regcheck.app) has received a Commendation from the Society for Improving Psychological Science! You can check the other awardees below too:
RegCheck
RegCheck is an AI tool to compare preregistrations with papers instantly.
regcheck.app
🏆We are very pleased to announce the 2026 Awards Winners and Commendations recipients for the Society for the Improvement of Psychological Science! 👏Join us in celebrating the exceptional projects that have revolutionized the field of psychological science. buff.ly/2lUKGFd
Research on cognitive offloading × AI is booming and getting lots of public attention. Unfortunately, some (or even many) of it raises serious credibility concerns. Like this paper, which others and I recently commented on via @pubpeer.com: pubpeer.com/publications...
PubPeer - AI Tools in Society: Impacts on Cognitive Offloading and the...
There are comments on PubPeer for publication: AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking (2025)
pubpeer.com
The methods section is the most important section of a paper. Without knowing the methods, the results are literally meaningless. So it's not OK to relegate the methods to the end. Because this suggests they are an optional afterthought rather than the essential core of the entire paper.
As long as a paper's methods are described *somewhere* I don't think it matters where they are located in a manuscript
In sum, skip the Atlantic on this one and read Monstrous Regiment, a great standalone Pratchett and a perfect book for right now, because it takes on propaganda, war, gender, and how religious zealots starve nations and ruin everything. Quote me on that, why don’t you.
"If I hated to be the bearer of bad news, I wouldn’t have specialized in research methods."
New blog post! Let's say an effect you are interested in varies between countries and you want to explain some of that cross-cultural variation. What could possibly go wrong? www.the100.ci/2026/05/20/f...
Austin et al. (2026) advocated for AI in hypothesis-driven scale validation research. In this preprint, led by @wchiggins.bsky.social, we (@malte.the100.ci @bethclarke.bsky.social) sketch out several weaknesses in this paper, and offer recommendations for avoiding them. osf.io/preprints/ps... 1/🧵
OSF
osf.io
New blog post! Let's say an effect you are interested in varies between countries and you want to explain some of that cross-cultural variation. What could possibly go wrong? www.the100.ci/2026/05/20/f...
Friends don’t let friends run moderated cross-country regressions
Header image: Moderate cross-country slopes (Photo: By Erik W. Kolstad - Flickr, CC BY 2.0,) Here’s a particular genre of article that relies on multiple data collection sites, such as studies anal...
the100.ci
There's an interesting conversation unfolding on PubPeer. Inconsistencies were flagged in a paper by a "global keynote speaker on AI&neuroscience" and then the author casually mentions on LinkedIn that she also has pre-post-fMRI for 300 people 👀> pubpeer.com/publications...
PubPeer - Generative artificial intelligence reliance and executive fu...
There are comments on PubPeer for publication: Generative artificial intelligence reliance and executive function attenuation: Behavioral evidence of cognitive offload in high-use adults (2026)
pubpeer.com
When patients ask, “What disorder do I really ‘have’?” the honest answer is usually more interesting and messier than a single diagnostic label. I wrote for the @nytimes.com on what I wish people understood about diagnoses and the nature of mental health problems. www.nytimes.com/2026/05/11/o...
Opinion | We’re Thinking About Mental Health Diagnoses All Wrong
nytimes.com
Hey psychologists, when you're explaining some psychological phenomenon that's perfectly well captured in psychological terms, maybe you don't need to refer to some shitty neuroimaging study with ten people to make it seem more "sciencey"...
This...the fetishizing of unreplicable neuroscience in order to legitimize our work is self-defeating. We don't need to know that depression is in the brain (as opposed to what???) to know it is problematic. How much money have we wasted proving something like depression is "real"?
Hey psychologists, when you're explaining some psychological phenomenon that's perfectly well captured in psychological terms, maybe you don't need to refer to some shitty neuroimaging study with ten people to make it seem more "sciencey"...
@taymalsalti.bsky.social has done trojan work trying to reproduce the results of three meta-analyses, finding the usual host of issues of effect sizes that are apparently incorrect or non-reproducible Our article is out now: doi.org/10.1111/ejn....
Ian Hussey (@ianhussey.mmmdata.io) contacted me about the recurring problem of unusually large SMD values appearing in published meta-analyses. Very large SMDs can result from using SEs in place of SDs in their calculation, which is a mistake I have also seen in practice. 1/4
Eagle-eyed @grinschglsandra.bsky.social spotted this mess of a plot, with bars that don't match percentages. We, with @malte.the100.ci, quantify a host of other issues we spotted here. pubpeer.com/publications...
If you are interested in EMDR, you probably heard the story about Francine Shapiro "discovering" the method during a walk in the park. This story is largely wrong.
A Photograph Exposes EMDR’s True Origins | Skeptical Inquirer
Sometimes if you keep searching, digging, and sifting, you’ll find a nugget that is bigger than anything you expected. That happened to us when we looked fo ...
skepticalinquirer.org
You ever try the 'Reading the Mind in the Eyes test' and wonder, how is that any of them?!
📣 PsychLing-101 — Deadline May 1 Join us in building a large-scale, trial-level psycholinguistic database for cumulative science and LLM evaluation. So far: 36 submissions and 40M+ observations All contributors become coauthors 🤝 🔗 Visit github.com/Data-X01/Psy...
GitHub - Data-X01/PsychLing-101: Large-scale, trial-level psycholinguistic database for cumulative science and LLM evaluation, driven by the community.
Large-scale, trial-level psycholinguistic database for cumulative science and LLM evaluation, driven by the community. - Data-X01/PsychLing-101
github.com
All the more reason to use systematic software designed to tackle this very problem, like RegCheck (regcheck.app)
RegCheck
RegCheck is an AI tool to compare preregistrations with papers instantly.
regcheck.app
Journal editors - the status quo on preregistration is not working! You need to check submissions vs. preregistrations before sending articles out for review. *56%* of experiments I reviewed in last year have severe problems with non-disclosure, undocumented deviations, & more - see Claude summary ↓
I’m hiring a PhD student! The candidate will work alongside @zefreeman.bsky.social, who is joining our research group as postdoc. jobs.unibe.ch/job-vacancie...
PhD Student in Meta-Science and Clinical Psychology - Universität Bern
Universität Bern is looking for PhD Student in Meta-Science and Clinical Psychology
jobs.unibe.ch
Good blog post that perfectly voices the frustration early career researchers can have when they hear senior researchers discuss preregistration. The continuing lack of understanding among senior researchers who keep voicing the same misunderstandings *is* frustrating and embarrassing.
Recently, I got quite frustrated listening to a large group of economists discussing preregistration and pre analysis plans. So, I've channelled that into an argument for why they should be preregistering their work whenever performing confirmatory research kdoroc.substack.com/p/arguing-wi...
I wrote about how I teach statistics. As I redesign for the AI era, I won't forget the benefits of multimodal, tangible representations. The Five Gs: Greek, Graphs, Grammar, Gadgets, and Games. In the new journal, Teaching Educational Research Methods: doi.org/10.5149/term...
New post on The 100% CI: Science needs downvotes. www.the100.ci/2026/04/13/s... In which I make the case that grant funders should add funding lines that include a module for bug bounties.
Science needs downvotes
A bug bounty module in grants would give criticism a leg up The Soviet Union was good at producing shoes. Factories made 800 million pairs a year, twice as many as Italy, three times as many as the...
the100.ci
For a short (German) summary of SCORE, check out @briannosek.bsky.social and my interview for @deutschlandfunk.de.web.brid.gy www.deutschlandfunk.de/glaubwuerdig...
Glaubwürdige Forschung: Sozialwissenschaftliche Studien oft nicht replizierbar
deutschlandfunk.de
SCORE, a collaboration of 865 researchers, is now released as three papers in Nature, six preprints, and a lot of data (cos.io/score/). SCORE examined repeatability of findings from the social-behavioral sciences and tested whether human and automated methods could predict replicability.
Very pleasantly surprised to see my recent preprint being discussed in this NYT article on silicon sampling and its implications for opinion polling. More forthcoming work on this coming soon! (H/T @rohanalexander.bsky.social) www.nytimes.com/2026/04/06/o...
Opinion | It’s Called Silicon Sampling, and It’s Going to Ruin Public Opinion Polling
nytimes.com
Depluralize a movie: The king‘s soliloquy
Depluralize a movie: Of a mouse and a man
Fun facts - The original study reported 23 IQ point gains - Jordan Peterson tweeted about it (sort of) - @jamiecummins.bsky.social new preregistered RCT shows null effects
A failure to find effects of relational operant training on scholastic aptitude of school children: A randomised controlled trial: https://osf.io/cauvz