Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...
Jamie Cummins
@jamiecummins.bsky.social
Advanced postdoc at University of Bern. Also sometimes @bennettoxford.bsky.social. Metascience stuff. Creator and developer of RegCheck (https://regcheck.app). STAR Editor at Psych Science. @error.reviews 🇮🇪
However if the goal isn't retention but is instead behaviour change, delay duration may be a useful tool among users who remain engaged. Our findings suggest that different design principles might optimise different intervention outcomes. Read the preprint here: osf.io/preprints/ps...
OSF
osf.io
Are you interested in smartphone interventions?📱 We are too. Specifically we were interested in increasing retention in a digital friction intervention. Let me tell you about our new preprint (1/n) with @dgruning.bsky.social @frederik.riedel.wtf @malte.the100.ci using one sec
I think this is kinda sorta how we did it so I can RT it.
Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...
Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...
A guide to evaluating LLM-based extraction and judgement tools: https://osf.io/2f4sk
In the last 2 years I (and our lovely collaborators) have been rather obsessively working on Metacheck - a tool to automatically check best practices in papers. In this RIOTS summerschool talk, I walk you through what the tool can do, and future plans www.youtube.com/watch?v=OWd2...
Metacheck: automatically check papers for best practices - Daniel Lakens
Can automated tools help researchers catch problems before their work is published? In this talk from the King’s Open Research Summer School 2026, Professor Daniël Lakens (Eindhoven University of…
youtube.com
Psychological Science is running a special issue on "Advancing Psychological Science using Large Language Models", co-edited by me and @ruben.the100.ci. Check out the calls for submissions below! www.psychologicalscience.org/publications...
Special Issue Call for Submissions: “Advancing Psychological Science using Large Language Models”
Editors: Jamie Cummins and Ruben C. ArslanPsychological Science invites submissions for a special issue on psychological science involving large language models. We welcome theoretically and empirical...
psychologicalscience.org
The Bennett conference returns to Oxford! December 8th and 9th Open code in science Read more and register www.bennett.ox.ac.uk/blog/2026/07...
Coming 8-9 December: The 2026 Bennett Oxford Symposium on Open Code in Science | Bennett Institute for Applied Data Science
Join us for the 2026 Bennett Oxford Symposium on Open Code in Science
bennett.ox.ac.uk
The Bennett Institute is having its annual conference - this year’s theme is on Open Code in Science. Come join us in Oxford on 8-9 December!
Registration is open for the 2026 Bennett Oxford Symposium on Open Code in Science! 📅 8-9 December 2026 📍 The Examination Schools, University of Oxford, 75-81 High St, Oxford OX1 4BG ℹ️ https://www.bennett.ox.ac.uk/events/2026-bennett-institute-symposium/ We look forward to seeing you there!
Registration is open for the 2026 Bennett Oxford Symposium on Open Code in Science! 📅 8-9 December 2026 📍 The Examination Schools, University of Oxford, 75-81 High St, Oxford OX1 4BG ℹ️ https://www.bennett.ox.ac.uk/events/2026-bennett-institute-symposium/ We look forward to seeing you there!
New blog post! Odds ratios are...a bit weird. Here I explain two problems (interpretation, non-collapsibility) and provide some simple recommendations for improved reporting. www.the100.ci/2026/07/30/r...
Reviewer notes: Odds ratios are really odd
When psychological researchers investigate binary outcomes, they routinely report odds ratios to quantify the strength of an effect or an association. This is quite understandable given that psycholog...
the100.ci
Synthetic survey participants cannot substitute for sample diversity in policy: https://osf.io/8x37h
Likewise and very good preview of more to come on this topic soon!
New paper (led by Xiangji Ying) showing LLMs can achieve high accuracy in detecting changes in outcomes across registrations versions. Thrilled to have been a small part of this great work. Check it out below!
We show that LLMs are good at detecting outcome definition changes between registration versions: An automated approach to improve clinical trial registration and to identify outcome changes on ClinicalTrials.gov rdcu.be/fwHWo @ndevito1.bsky.social @mjpages.bsky.social @jamiecummins.bsky.social
An automated approach to improve clinical trial registration and to identify outcome changes on ClinicalTrials.gov
npj Digital Medicine - An automated approach to improve clinical trial registration and to identify outcome changes on ClinicalTrials.gov
rdcu.be
New paper (led by Xiangji Ying) showing LLMs can achieve high accuracy in detecting changes in outcomes across registrations versions. Thrilled to have been a small part of this great work. Check it out below!
We show that LLMs are good at detecting outcome definition changes between registration versions: An automated approach to improve clinical trial registration and to identify outcome changes on ClinicalTrials.gov rdcu.be/fwHWo @ndevito1.bsky.social @mjpages.bsky.social @jamiecummins.bsky.social
New preprint w/ Leo Tiokhin & @lakens.bsky.social! Registered Reports (RRs) are great for science b/c publication is guaranteed before results are known, reducing publication bias & QRPs. This feature supposedly also makes them great for *scientists*. But is that really true? 1/ tinyurl.com/4sy993rr
Career incentives make Registered Reports unattractive under realistic academic conditions
Registered Reports are a publication format designed to reduce publication bias by guaranteeing publication before results are known. This guarantee is considered attractive in the prevailing ‘publish...
tinyurl.com
RegCheck ( regcheck.app) has received a Commendation from the Society for Improving Psychological Science! You can check the other awardees below too:
RegCheck
RegCheck is an AI tool to compare preregistrations with papers instantly.
regcheck.app
🏆We are very pleased to announce the 2026 Awards Winners and Commendations recipients for the Society for the Improvement of Psychological Science! 👏Join us in celebrating the exceptional projects that have revolutionized the field of psychological science. buff.ly/2lUKGFd
🏆We are very pleased to announce the 2026 Awards Winners and Commendations recipients for the Society for the Improvement of Psychological Science! 👏Join us in celebrating the exceptional projects that have revolutionized the field of psychological science. buff.ly/2lUKGFd
SIPS 2026 Awards Announced!
We are very pleased to announce the 2026 Awards Winners for the Society for the Improvement of Psychological Science! Join us in celebrating the exceptional projects that have revolutionized the fi…
improvingpsych.org
RegCheck is going straight into the editorial triage workflow for testing at EBT www.tandfonline.com/journals/teb...
Evidence-Based Toxicology
An open science journal publishing evidence-based research in the toxicological and human environmental health sciences, investigating how environmental exposur
tandfonline.com
Comparing registrations to published papers is essential to research integrity - and almost no one does it routinely because it's slow, messy, and time-demanding. RegCheck was built to help make this process easier. Today, we launch RegCheck V2. 🧵 regcheck.app
There are still a few days to apply to work with me and @malte.the100.ci as a PhD or postdoc on the development and evaluation of RegCheck. Help us research whether RegCheck works in practice in helping to reduce preregistration-paper discrepancies. Apply below!
WE ARE HIRING! Come work with us (me + @malte.the100.ci) on developing and validating RegCheck. PhD or postdoc applications welcome. Folk with backgrounds in meta-science, psychology, (pre)clinical trials, and/or economics are especially invited to apply. jobs.unibe.ch/job-vacancie...
New paper out now 🥳 When psychologists discuss generalisability, they often refer to vague notions of representativeness. We provide an accessible intro to the total survey error framework as a tool to reason about this more rigorously. w @taymalsalti.bsky.social @ruben.the100.ci >
Since some have asked 'Where does the logo for the Clinical Trials Abundance blog come from?' The answer is sadly not very deep or sophisticated. It comes from the 'moar' meme (knowyourmeme.com/memes/moar) and started as a placeholder logo before we came up with a better one (which we didn't).
If you had to guess, what proportion of social/personality psych articles include a claim that the findings can be applied in the real world?
3. Additional discrepancies detected with the RegCheck tool (🙏to @jamiecummins.bsky.social) included the assessment periods (4 weeks and 3 months vs. days 30 and 60), or that the psychiatrist-confirmed diagnosis requirement specified in the protocol is missing from the publication.
1. We wrote a letter to the editor about this fluoxetine trial for children aged 4-6. Given the urgency and the expected delay in response, we also published it as a preprint, available here: doi.org/10.31234/osf... Which problems did we detect with this study ? 🧵
Here's the first trial to test fluoxetine for #depression in children under 6: psych-partners.com/fluoxetine-f... Positive signal, but agitation is a risk, and a single study doesn't change the fact that Parent-Child Interaction #Therapy has the best evidence in this age group. #psychiatry #pmhnp
Today we launch the first stable release of RegCheck: v1.0.0. RegCheck makes it easier and quicker to compare study registrations to published papers for consistency - something we know is important in principle, but rarely done in practice. A 🧵 on what's new: regcheck.app
RegCheck
RegCheck is an AI tool to compare preregistrations with papers instantly.
regcheck.app
This is amazing and I hippo it gets used widely.
Today we launch the first stable release of RegCheck: v1.0.0. RegCheck makes it easier and quicker to compare study registrations to published papers for consistency - something we know is important in principle, but rarely done in practice. A 🧵 on what's new: regcheck.app
Fantastic use of AI.
Today we launch the first stable release of RegCheck: v1.0.0. RegCheck makes it easier and quicker to compare study registrations to published papers for consistency - something we know is important in principle, but rarely done in practice. A 🧵 on what's new: regcheck.app