Gordon Forbes

@gforb.bsky.social

I love Netflix for their data science blog and The BBC for their ggplot2 resources.

Did you spend 4-6 weeks writing your last analysis plan then get mocked by statisticians on the internet for being slow 😜. We had a go doing it with AI - still takes 4-6 weeks but the AI will do the boring parts letting you conentrate on the stats. Pre print-here bsky.app/profile/gfor...

Darren Dahly@statsepi.bsky.social · 6mo ago

I'd be curious to know why it takes 4-6 weeks to develop a SAP if you have the protocol and an existing wealth of experience developing SAPs from protocols.

Hey #rstats, What's your rule for splitting R scripts that form part of a wider analysis pipeline / project? I usually write a single script which includes sections for each step from data cleaning to the final results, but it can become unwieldy when the script becomes long. ...

On tabular health data, time and time again, I see linear (or generalised linear models) perform as well or better than machine learning algorithms that avoid linearity assumptions. I am surprised by this, as the linearity assumption is unlikely to be true. Does anyone else see this? Why is this?

The optimal machine learning model was linear regression

Happy to see that ordered beta regression reached 100 citations on Google Scholar! The model has citations from work in climate science, ecology, medicine, psychology, & political science, just to name a few. Thanks to all of you for using ordbetareg (or glmmTMB)!! #rstats

Bild

The term "digital twin" as it is now used in medicine has no real relationship to how the term is used in engineering. Yet every paper on the former talks about the success of the latter as if that's relevant. They are not the same!

My spell check is trying to change bootstrap to Boomer. As in Boomer p-values. I didn't know the generation wars had made it to statistical inference. What next? Millennial credible intervals?

A minor stylistic preference I’ve recently found myself using: When introducing a key initialism or acronym in a paper, put the compressed version in the text and its expansion in parentheses. Instead of ‘under missing at random (MAR)’, use ‘under MAR (missing at random)’. 1/

What advice do folks have for organising projects that will be deployed to production? How do you organise your directories? What do you do if you're deploying multiple "things" (e.g. an app and an api) from the same project?

Am looking for a particular article that argued that results of "explanatory" trials might generally be more timeless than those of pragmatic trials, since the latter might be reliant on a particular context at a given moment in time. Can't remember who this was by - something from imp science lit?