Greg Landrum

@greglandrum.bsky.social

Cheminformatician, developer, climber, runner, hiker, cook ORCID: 0000-0001-6279-4481

Our new perspective in @jacs.acspublications.org makes some recommendations for building ML models for reaction outcomes. We didn't want to go too far out on a limb and call them "best practices" ;-) pubs.acs.org/doi/10.1021/...

Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes

Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.

pubs.acs.org

I listened to a couple of 10,000 Maniacs albums last week and then had bits from their songs looping through my head for a week. I'm finally over that, but Spotify is really trying hard to infect me again.

One of my favorite parts of the dystopia that we are now heading into has got to be waiting for the "are you a human" challenges to resolve on EVERY SINGLE WEB PAGE that has content that's worth seeing Thank you AI companies for unapologetically stealing all the things!

We did a new major #RDKit release yesterday. v2026.03.1 is out and the conda-forge builds are available. As always, especially with major releases, be sure to read the "Backwards incompatible changes" section of the release notes! github.com/rdkit/rdkit/...

Release 2026_03_1 (Q1 2026) Release · rdkit/rdkit

Release_2026.03.1 (Changes relative to Release_2025.09.1) Acknowledgements (Note: I'm no longer attempting to manually curate names. If you would like to see your contribution acknowledged with you...

github.com

More terrible UI, this time from the @github.com app. I'm in landscape mode, vertical space is the limiter, but look at the incredible amount that is being wasted. Never mind that much of the info doesn't need to be displayed on that screen.

Screenshot from the github app where more than half the vertical space is taken up with UI elements and not the code I’m trying to view