David Selby

@davidselby.bsky.social

Data science researcher working on applications of machine learning in health at DFKI, getting the most out of small data. Reproducible #Rstats evangelist and unofficial British cultural ambassador to Rhineland-Palatinate 🇩🇪 https://selbydavid.com

📊📉📈 Better data visualizations with AI: can LLMs provide constructive critiques on existing charts? We explore how generative AI can automate #MakeoverMonday -type exercises, suggesting improvements to existing charts. 📄 New preprint + benchmark dataset 💽 arxiv.org/abs/2508.05637

Automated Visualization Makeovers with LLMs

Making a good graphic that accurately and efficiently conveys the desired message to the audience is both an art and a science, typically not taught in the data science curriculum. Visualisation makeo...

arxiv.org

What is a "Visible Neural Network"? It's a new kind of deep learning model for multi-omics, where prior knowledge and interpretability are baked into the architecture. 📄 We reviewed dozens of models, datasets & applications, and call for better tools/benchmarks: www.frontiersin.org/journals/art...

Frontiers | Visible neural networks for multi-omics integration: a critical review

BackgroundBiomarker discovery and drug response prediction are central to personalized medicine, driving demand for predictive models that also offer biologi...

frontiersin.org

Just published: 'Had enough of experts? Quantitative retrieval from large language models' Can LLMs, having read the scientific literature, offer us useful numerical info to help fill in missing data and fit statistical models, like a real human expert? We investigate: doi.org/10.1002/sta4...

Had Enough of Experts? Quantitative Knowledge Retrieval From Large Language Models

Large language models (LLMs) have been extensively studied for their ability to generate convincing natural language sequences; however, their utility for quantitative information retrieval is less w...

doi.org

Paper just accepted in Stat! Can LLMs replace experts as sources of numerical information, such as Bayesian prior distributions for statistical models, or filling in missing values in tabular datasets for ML tasks? We evaluate on applications across different fields. arxiv.org/abs/2402.07770

Had enough of experts? Quantitative knowledge retrieval from large language models

Large language models (LLMs) have been extensively studied for their abilities to generate convincing natural language sequences, however their utility for quantitative information retrieval is less w...

arxiv.org