Alexander Gibson

@alexdgibson.bsky.social

metascientist www.alexdgibson.com

Outstanding work by @alexdgibson.bsky.social to uncover the use of highly dubious data sets from @kaggle.com being used in hundreds of research papers and potentially even informing clinical practice. If you're re-using data, take the time to confirm that it's real. www.nature.com/articles/d41....

Dozens of AI disease-prediction models were trained on dubious data

The models are designed to predict someone’s risk of diabetes or stroke. A few might already have been used on patients.

nature.com

New Blog: Learning R for Good Research Practices‼️ Read Part 1: alexdgibson.com/blog/r_1/ This is the first of a multipart series where I go into my experiences learning R, highlighting tips and resources that were essential in my learning.

Learning R for Good Research Practices: Part 1 | Alexander Gibson

In this multi-part blog series I outline some steps to help start learning R. Part 1 is a background on R, packages and resources and tips on how to get started learning.

alexdgibson.com

We need more people reading more books! Such an important skill and has done so much for me. There’s nothing better than a good book, a change in perspective, beliefs or the generation of new ideas.

Bild

Ok, time for a short thread about this paper. My sense over the past six months or so is that chain-of-thought prompting as used in e.g. ChatGPT o.3 improves substantially upon previous systems such as ChatGPT 4.o, at least for certain tasks. But how revolutionary is it?

Carl T. Bergstrom@carlbergstrom.com · last yr.

If I have time I'll put together a more detailed thread tomorrow, but for now, I think this new paper about limitations of Chain-of-Thought models could be quite important. Worth a look if you're interested in these sorts of things. ml-site.cdn-apple.com/papers/the-i...

The Illusion of Thinking:
Understanding the Strengths and Limitations of Reasoning Models
via the Lens of Problem Complexity
Parshin Shojaee∗† Iman Mirzadeh∗ Keivan Alizadeh
Maxwell Horton Samy Bengio Mehrdad Farajtabar
Apple
Abstract
Recent generations of frontier language models have introduced Large Reasoning Models
(LRMs) that generate detailed thinking processes before providing answers. While these models
demonstrate improved performance on reasoning benchmarks, their fundamental capabilities, scal-
ing properties, and limitations remain insufficiently understood. Current evaluations primarily fo-
cus on established mathematical and coding benchmarks, emphasizing final answer accuracy. How-
ever, this evaluation paradigm often suffers from data contamination and does not provide insights
into the reasoning traces’ structure and quality. In this work, we systematically investigate these
gaps with the help of controllable puzzle environments that allow precise manipulation of composi-
tional complexity while maintaining consistent logical structures. This setup enables the analysis
of not only final answers but also the internal reasoning traces, offering insights into how LRMs
“think”. Through extensive experimentation across diverse puzzles, we show that frontier LRMs
face a complete accuracy collapse beyond certain complexities. Moreover, they exhibit a counter-
intuitive scaling limit: their reasoning effort increases with problem complexity up to a point, then
declines despite having an adequate token budget. By comparing LRMs with their standard LLM
counterparts under equivalent inference compute, we identify three performance regimes: (1) low-
complexity tasks where standard models surprisingly outperform LRMs, (2) medium-complexity
tasks where additional thinking in LRMs demonstrates advantage, and (3) high-complexity tasks
where both models experience complete collapse. We found that LRMs have limitations in exact
computation: they fail to use explicit …

Reading over my brother’s undergraduate assignment. Criteria has a section: “Is there an appropriate amount of detail in this section so that you could replicate this study?” I don’t remember learning about replication in my undergrad! Nice to see it getting some air time.

If you're currently working—or have worked in the past 5 years—in consulting or collaborative research as a biostatistician, we’d love to hear from you. 📅 Closes 11th July This survey has been approved by the QUT Human Research Ethics Committee (approval #9691).