Barnabas Szaszi

@szaszibarnabas.bsky.social

Behavior change at scale, decision-making, meta-science, inequality, and a 'little bit' of happiness. Fulbright Alumn at Harvard & Columbia. Currently leading the Behavioral Science Center in Budapest.

New blog post: Evaluating Dr. Cuddy’s Claim that the Debunking of Power Posing is a Myth. daniellakens.blogspot.com/2026/05/eval... On an AI generated description of a non-existent study, incorrectly citing findings from studies, and the importance of scientific criticism.

Evaluating Dr. Cuddy’s Claim that the Debunking of Power Posing is a Myth

In this blog post I will analyse the arguments that Dr. Amy Cuddy provided in a blog post “The "Power Posing Was Debunked" Myth: What the Re...

daniellakens.blogspot.com

Per protocol analysis strikes again! Folks, if you randomize but then don‘t analyze some of the people who got randomized (maybe because they didn’t adhere to instructions, maybe because they dropped out), randomization will no longer do all the heavy causal inference lifting.

Peter Tennant@pwgtennant.bsky.social · 6mo ago

This recent RCT of an "AI stethoscope" claims the technology "shows promise" for diagnosing cardiovascular conditions. It does not. It is a textbook example of the risks of conducting unprincipled 'per protocol analyses'. Once again, peer review at a major medical journal has failed. 🧵 1/

Our lab has the capacity to test ~500 uni students each semester If you’re a researcher in cognitive psychology or metascience and need data collection support, we’d love to collaborate. We can help collect high-quality data from a large student sample. Get in touch to discuss potential projects!

It's ironic to see a discipline care **so much** about unbiasedness (causal inference!) at the level of a single test but then have a research production system and culture that is basically a ferocious bias generation machine. This is not good.

The Iowa Gambling Task is an extreme example of Jingle Fallacy and schmeasurement. In 100 articles we found 244 different ways of scoring it, 177 were never reused. Correlations between them range -.99 to .99. At the same time, we show meta-analyses combine these results as if they’re equivalent.

Annika Külpmann@anniria.bsky.social · 7mo ago

How many versions of the Iowa Gambling Task (IGT) exist? And how much does this affect research using the IGT? More than you might think. 🧵

After 5 years of data collection, our WARN-D machine learning competition to forecast depression onset is now LIVE! We hope many of you will participate—we have incredibly rich data. If you share a single thing of my lab this year, please make it this competition. eiko-fried.com/warn-d-machi...

WARN-D machine learning competition is live » Eiko Fried

If you share one single thing of our team in 2026—on social media or per email with your colleagues—please let it be this machine learning competition. It was half a decade of work to get here, especi...

eiko-fried.com

Do you want to work with me?:) Please spread the word! We are looking for talented Post-doc candidates for a 10-month Junior Fellowship at the Behavioral Science Center, hosted by the Corvinus Institute for Advanced Studies (CIAS). 1/6 www.the-bs-lab.com

Behavioral Science

What do we do? We conduct large-scale behavioral science studies to improve the daily decisions, behavior, and experience of vulnerable individuals (e.g., the well-being of citizens and families / ed...

the-bs-lab.com

There still seems to be a lot of confusion about significance testing in psych. No, p-values *don’t* become useless at large N. This flawed point also used to be framed as "too much power". But power isn't the problem – it's 1) unbalanced error rates and 2) the (lack of a) SESOI. 1/ >

Post nicht verfügbar.

Thank you for sharing -- I was unaware of this framework for choosing among replication targets. I think I stand with the critics. Of course I'm left still not knowing how to choose among empirical estimands! Separately, this figure is amazing:

a good figure that skewers an "anonymous" reviewer for requiring many citations to (presumably) their work

Can large language models stand in for human participants? Many social scientists seem to think so, and are already using "silicon samples" in research. One problem: depending on the analytic decisions made, you can basically get these samples to show any effect you want. THREAD 🧵

The threat of analytic flexibility in using large language models to simulate human data: A call to attention

Social scientists are now using large language models to create "silicon samples" - synthetic datasets intended to stand in for human respondents, aimed at revolutionising human subjects research. How...

arxiv.org

HELP! 🙏 I’m looking for validated emotionally neutral texts (ideally pretested or normed) to use in a control condition. Participants will read the text and count letters. Any leads, references, or resources? 🙏

PCI Psychology is here!! 🎉🥳 After over a year of hard work by so many people, we are thrilled to announce that we are open for submissions! Join us in making publishing more efficient, equitable, and open: psych.peercommunityin.org #PsychSciSky #scipub

PCI Psychology

Peer Community in Psychology

psych.peercommunityin.org

PsyArXivBot@psyarxivbot.bsky.social · last yr.

Peer Community in Psychology: A platform for peer review of preprints across psychology: https://osf.io/m456e

A recently published meta-analysis in Nature Human Behaviour "found evidence supporting the efficacy of social comparison as a behaviour change technique in shaping behaviour in the desired direction". I was curious, so I re-analyzed the manuscript, but the funnel plots below say it all.

BildBild