"The Illusion of Comparability Among Standardised Effect Sizes: Why Education Evaluations Should Report Raw Effects" New @cgdev.org working paper (with Rossiter, Hares, and Henny) cgdev.org/publication/... Summary blog post cgdev.org/blog/standar...
Winston Lin
@linstonwin.bsky.social
senior lecturer in statistics, penn NYC & Philadelphia https://www.stat.berkeley.edu/~winston
When I reflect on what shaped me most as a scholar, one seminar stands out: Ruth Collier’s dissertation writing workshop at Berkeley. I wrote about its influence on my research, teaching, and scholarship in general. jaeyeonkim.substack.com/p/the-ruth-c...
The Ruth Collier Method
Think Clearly and Write Honestly
jaeyeonkim.substack.com
Good news everyone 🥳 Our (w @vincentab.bsky.social) primer on models as prediction machines (with the marginaleffects package) is finally officially published!> journals.sagepub.com/doi/10.1177/...
Models as Prediction Machines: How to Convert Confusing Coefficients Into Clear Quantities - Julia M. Rohrer, Vincent Arel-Bundock, 2026
Psychological researchers usually make sense of regression models by interpreting coefficient estimates directly. This works well enough for simple linear model...
journals.sagepub.com
Journal editors - the status quo on preregistration is not working! You need to check submissions vs. preregistrations before sending articles out for review. *56%* of experiments I reviewed in last year have severe problems with non-disclosure, undocumented deviations, & more - see Claude summary ↓
You guys @carlislerainey.bsky.social has a free textbook online and it seems really useful pos5747.github.io/notes/
Bilingual joke? French Wikipedia says the Poisson distribution is "not to be confused with Fisher's distribution" (the F-distribution) fr.wikipedia.org/wiki/Loi_de_...
Rosenbaum's Observation and Experiment is great too. I have sadly not read his more technical books yet. www.hup.harvard.edu/books/978067...
Observation and Experiment — Harvard University Press
A daily glass of wine prolongs life—yet alcohol can cause life-threatening cancer. Some say raising the minimum wage will decrease inequality while others say it increases unemployment. Scientists onc...
hup.harvard.edu
You need to bring in the same toolkit as in studies that try to establish causality without randomization. I know it sounds unfair, but I don’t make the rules. These situations are instances of post-treatment bias, if you want to read up on it as a psychologist:
Causal inference for psychologists who think that causal inference is not for them
Correlation does not imply causation and psychologists' causal inference training often focuses on the conclusion that therefore experiments are needed—without much consideration for the causal infer...
compass.onlinelibrary.wiley.com
A more user friendly t-test regression variable description frequency plots, and more. datacolada.org/132
🚨SOLUTIONS🚨 Desk reject more stuff with actionable feedback. Don’t request second reviews Build larger editorial boards of volunteers Wait to submit your work until it’s ready; a.k.a don’t send in your half-baked trash hoping for feedback 6/7
After years in academia, I’m exploring data science and research roles in industry. I'm a quant. social scientist (PhD Yale ’24, NYU) focused on causal inference, experiments, and large-scale data. Feel free to get in touch or share; all leads appreciated. dwstommes@gmail.com
This quote also reminds me of something that we wrote in our paper on path analysis (journals.sagepub.com/doi/10.1177/...). People are just expecting *way* too much of a single study, literally new discoveries exceeding Gregor Mendel's.
Comparing registrations to published papers is essential to research integrity - and almost no one does it routinely because it's slow, messy, and time-demanding. RegCheck was built to help make this process easier. Today, we launch RegCheck V2. 🧵 regcheck.app
RegCheck
RegCheck is an AI tool to compare preregistrations with papers instantly.
regcheck.app
“Coding for humans: Best practices for writing software people can read” statmodeling.stat.columbia.edu/2026/01/17/c...
“Coding for humans: Best practices for writing software people can read” | Statistical Modeling, Causal Inference, and Social Science
statmodeling.stat.columbia.edu
Accessibility is *absolutely* key but also hard because of the curse of knowledge. I've written down some writing advice here: www.the100.ci/2024/12/01/w.... If you're more of a technical person, consider teaming up with a substantive researcher for instant audience access.>
Writing about technical topics in an accessible manner
A wise man – I’m quite sure it was Brian Wansink – once pointed out that it is impossible to both read and write a lot. So, maybe reading a post about how to write just steals time from the more urgen...
the100.ci
Some people bring up (1) the cost of criticism and (2) that a lot of criticism has already been voiced but ignored. Both points are valid, so here are some suggestion for (1) reducing backlash and (2) increasing impact (from this talk of mine: juliarohrer.com/wp-content/u...
CIIG Seminar: Julia Rohrer | Making Rigorous Causal Inference More Mainstream | 20 Oct 2025
YouTube video by the Causal Inference Interest Group (CIIG)
youtube.com
Here's a suggestion for a New Year's resolution: If you see influential bad research, say something. One part of the whole replication crisis story is that a lot of psychological researchers privately knew that a lot of stuff was bad, but it wasn't discussed publicly.
Some closing thoughts for my students this semester on LLMs and learning #rstats datavizf25.classes.andrewheiss.com/news/2025-12...
Gentle reminder that a correlation coefficient isn’t a particularly great way to quantify the effect of a dichotomous treatment. See also www.the100.ci/2025/07/28/w...
What’s in a correlation?
Correlation may not imply causation, but let’s just ignore that for a second. Correlations are standardized effect size metrics and as such have some quirks by design. These are benign enough when you...
the100.ci
Most scientists don't understand how effect sizes work and are therefore far too quick to dismiss "small" effects. A correlation of .03 between taking aspirin & prevention of future heart attacks implied the prevention of 85 attacks in a sample of 10,845 people journals.sagepub.com/doi/10.1177/...
Excellent new editorial and guideline on interpreting p values and interval estimates bjsm.bmj.com/content/earl...
Interpreting p values and interval estimates based on practical relevance: guidance for the sports medicine clinician
Statistical methods are employed in medical research to estimate effects of treatments or health conditions across populations.1 2 This paper presents a framework to avoid common misinterpretations th...
bjsm.bmj.com
Are you or one of your students considering doing a Ph.D. in a social science? I've spent a lot of time talking about this w/ students & finally wrote something up. IMO, there are only 3 good reasons to do it. One of them needs to be true--otherwise, don't. medium.com/the-quantast...
The Only Three Reasons to Do a Ph.D. in the Social Sciences
If none are true, don’t do it.
medium.com
A nice recent article on why you should abandon hazard ratios. #statssky #episky
How hazard ratios can mislead and why it matters in practice - European Journal of Epidemiology
Hazard ratios are routinely reported as effect measures in clinical trials and observational studies. However, many methodological works have raised concerns about the interpretation of hazard ratios ...
doi.org
See our No-Spin report on a widely-covered NBER study of Medicaid expansion. In brief: Despite the abstract's claims that expansion reduced adult mortality 2.5%, the study found much smaller effects that fell short of statistical significance in its main preregistered analysis.🧵
Starting to look like I might not be able to work at Harvard anymore due to recent funding cuts. If you know of any open statistical consulting positions that support remote work or are NYC-based, please reach out! 😅
Issues with interpreting p-values haunts even AI, which is prone to same biases as human researchers. ChatGPT, Gemini & Claude all fall prey to "dichotomania" - treating p=0.049 & p=0.051 as categorically different, and paying too much attention to significance. www.cambridge.org/core/journal...
NEW: CONSORT 2025 now published! Some notable changes: -items on analysis populations, missing data methods, and sensitivity analyses -reporting of non-adherence and concomitant care -reporting of changes to any study methods, not just outcomes -and lots of other things www.bmj.com/content/389/...
CONSORT 2025 explanation and elaboration: updated guideline for reporting randomised trials
Critical appraisal of the quality of randomised trials is possible only if their design, conduct, analysis, and results are completely and accurately reported. Without transparent reporting of the met...
bmj.com
How to write a response to reviewers. www.sciencedirect.com/science/arti...
How To Write a Response to Reviewers
sciencedirect.com
Very nice explainer by @economictricks.bsky.social www.econometrics.blog/post/why-eco...
Why Econometrics is Confusing Part 1: The Error Term | econometrics.blog
“Suppose that \(Y = \alpha + \beta X + U\).” A sentence like this is bound to come up dozens of times in an introductory econometrics course, but if I had my way it would be stamped out completely.
econometrics.blog
today we will all read imbens 2021 on statistical significance and p values, which is a strong contender for having the best opening paragraph of any stats paper pubs.aeaweb.org/doi/pdf/10.1...
Here's some older, related stuff (from me) aimed at political scientists. Related paper #1 "Arguing for a Negligible Effect" Journal: onlinelibrary.wiley.... PDF: www.carlislerainey.c...