Maarten van Smeden
@maartenvsmeden.bsky.social
statistician • associate prof • team lead health data science and head methods research program at julius center • director ai methods lab, umc utrecht, netherlands • views and opinions my own
Lang gewerkt aan de reconstructie van dit verhaal. Zaterdag in NRC, vandaag al online: Heeft de geit het weer gedaan? Het bewijs is broodmager dat omwonenden longontstekingen krijgen door de geitenhouderij Cadeaulink 🎁 www.nrc.nl/nieuws/2026/...
Heeft de geit het weer gedaan? Het bewijs is broodmager dat omwonenden longontstekingen krijgen door de geitenhouderij
Volksgezondheid: Zijn geiten opnieuw een gezondheidsrisico, na de eerdere uitbraak van Q-koorts? Zorgen ze voor longontstekingen? Onderzoek dat dat zou aantonen, wordt „volledig mislukt” genoemd.
nrc.nl
Normality : a very rare shape for data www.linkedin.com/posts/maarte...
Statistical terms: what they really mean Multicolinearity— these variables all look the same Heteroscedasticity— the variation varies Attenuation— being too modest Overfitting— too good to be… | Maar...
Statistical terms: what they really mean Multicolinearity— these variables all look the same Heteroscedasticity— the variation varies Attenuation— being too modest Overfitting— too good to be true Co...
linkedin.com
New post, on why we should stop telling people the normal distribution is in places that it isn't: kucharski.substack.com/p/no-the-nor...
Came across this gem again today. Still makes me giggle @richarddriley.bsky.social @gscollins.bsky.social Yes, it is published: pubmed.ncbi.nlm.nih.gov/32817260/
🎯So many important points packed into this editorial. ht @maartenvsmeden.bsky.social www.nature.com/articles/s41...
Show us the evidence for the value of medical AI - Nature Medicine
Claims that medical AI is improving care must be backed by appropriate evidence.
nature.com
Appraising ML prediction models @maartenvsmeden.bsky.social 1) Share model and code 2) Predictors and language 3) Class imbalance & calibration 4) Model validation sample size 5) Is the model clinically useful in practice? www.jclinepi.com/article/S089...
Preliminary appraisal of machine learning based prediction models
A significant portion of healthcare research is devoted to the development of prediction models, yet the integration of these models into routine clinical care remains limited. Persistent barriers, su...
jclinepi.com
Fantastic new paper casting doubt on explainability of explainable AI. To explain complex machine learning algorithms you need reproducibility of the explanation at a minimum academic.oup.com/ehjdh/articl... #machinelearning #Statistics #StatsSky @maartenvsmeden.bsky.social
Signal or noise? Evaluating commonly used attribution methods for explaining deep neural networks in electrocardiogram classification
AbstractAims. Attribution-based explainability methods are widely used in electrocardiogram (ECG) analysis to interpret predictions from ‘black-box’ deep n
academic.oup.com
Happy to see this in print! doi 10.1146/annurev-statistics-042324-123749 @maartenvsmeden.bsky.social @laurewynants.bsky.social @vanamsterdam.bsky.social and Ewout Steyerberg
I have got the data and I figured out exactly how to do my cluster analysis all I need is a relevant question that my cluster analysis is going to answer
What if you combine open datasets with AI? Apparently, a 3 fold increase in low quality research papers, mass-produced by paper mills. Interesting study in @jclinepi.bsky.social #academicsky #episky #medsky #Skystats Thanks to @maartenvsmeden.bsky.social for initially posting this on Linkedin!
Our guidance regarding performance measures for medical AI models is finally out! - Stop bashing AUROC, although it does not settle things - Calibration and clinical utility are key - Show risk distributions - Classification statistics (e.g. F1) are improper www.thelancet.com/journals/lan...
Evaluation of performance measures in predictive artificial intelligence models to support medical decisions: overview and guidance
Numerous measures have been proposed to illustrate the performance of predictive artificial intelligence (AI) models. Selecting appropriate performance measures is essential for predictive AI models i...
thelancet.com
NEW PAPER The use of explainable AI in healthcare evaluated using the well known Explain, Predict and Describe taxonomy by Galit Shmueli link.springer.com/article/10.1...
this is one of my favourite observations about sample size calculations. (afaik first articulated by Miettinen in 1985)
For some research studies the optimal sample size should be estimated at 0
For some research studies the optimal sample size should be estimated at 0
“Data available upon reasonable request” is academic language for you can get my data OVER MY DEAD BODY
Manuscript_Final_Version_actualFINALcopy_version9b_USETHISONE.docx
Kind reminder: data driven variable selection (e.g. forward/stepwise/univariable screening) makes things *worse* for most analytical goals
NEW FULLY FUNDED PHD POSITION Looking for a motivated PhD candidate to join our team. Together with Danya Muilwijk, Jeffrey Beekman and I, you will explore opportunities and limitations of AI in the context of organoids For more info and for applying 👉 www.careersatumcutrecht.com/vacancies/sc...
Vacancy — PhD position on AI methodology for prediction of patient outcomes using organoid models
Are you passionate about bringing personalized medicine to the next level and make real impact in healthcare? Join our team and develop novel AI methodology to improve predictions of relevant patient ...
careersatumcutrecht.com
Interpretable "AI" is just a distraction from safe and useful "AI"
I wonder who those people are who come here dying to know what GenAI has done with some prompt you put in
If you think AI is cool, wait until you learn about regression analysis
NEW PREPRINT Explainable AI refers to an extremely popular group of approaches that aim to open "black box" AI models. But what can we see when we open the black AI box? We use Galit Shmueli's framework (to describe, predict or explain) to evaluate arxiv.org/abs/2508.05753
The healthcare literature is filled with "risk factors". This word combination makes research findings sound important by implying causality, while avoiding direct claims of having identified causal associations that are easily critiqued.
When forced to make a choice, my choice will be logistic regression model over linear probability model 103% of the time
Post just up: Is multiple imputation making up information? tldr: no. Includes a cheeky simulation study to demonstrate the point. open.substack.com/pub/tpmorris...
You can have all the omni-omics data in the world and the bestest algorithms, but eventually a predicted probability is produced & it should be evaluated using well-established methods, and correctly implemented in the context of medical decision making. statsepi.substack.com/i/140315566/...