Reason for declining to review: "Conflict of interest" Comments to Editor: "I choose not to donate my time to for-profit publishers."
Greg Landrum
@greglandrum.bsky.social
Cheminformatician, developer, climber, runner, hiker, cook ORCID: 0000-0001-6279-4481
Our new perspective in @jacs.acspublications.org makes some recommendations for building ML models for reaction outcomes. We didn't want to go too far out on a limb and call them "best practices" ;-) pubs.acs.org/doi/10.1021/...
Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
pubs.acs.org
This was an interesting one to work on. #chemsky
Our recent JCIM publication looks at the balance of data quality and quantity when building ML models for bioactivity using ChEMBL data. Spoiler: the results are surprisingly insensitive to dataset size when using a realistic validation scenario. pubs.acs.org/doi/full/10....
@pubs.acs.org this is a real problem for those of us who use RSS readers to stay on top of what's appearing in your journals.
Have just realized that none of my ACS RSS feeds are updating in any of my RSS readers. But when viewed in a browser the feed is OK...only after passing the Cloudflare human detection page. @pubs.acs.org can you please have a look at this? I see that there are others reporting this too. #chemsky
Have just realized that none of my ACS RSS feeds are updating in any of my RSS readers. But when viewed in a browser the feed is OK...only after passing the Cloudflare human detection page. @pubs.acs.org can you please have a look at this? I see that there are others reporting this too. #chemsky
I listened to a couple of 10,000 Maniacs albums last week and then had bits from their songs looping through my head for a week. I'm finally over that, but Spotify is really trying hard to infect me again.
“You can now set your availability in the reviewer dashboard to pause review invitations.” WTF? Do the journals think we work for them or something?
TFW you realize that you cut the fruit for your breakfast cereal on the same cutting board you used to mince garlic the night before.
One of my favorite parts of the dystopia that we are now heading into has got to be waiting for the "are you a human" challenges to resolve on EVERY SINGLE WEB PAGE that has content that's worth seeing Thank you AI companies for unapologetically stealing all the things!
A "Yay! Private Equity!" thread.... I used to buy a license for MarvinSketch every year because it's a great tool and was pretty inexpensive. 1/n
Hey LinkedIn: 🖕for letting companies pay you to spam me. And 🖕🖕🖕for sending me reminders when I ignore that spam.
We did a new major #RDKit release yesterday. v2026.03.1 is out and the conda-forge builds are available. As always, especially with major releases, be sure to read the "Backwards incompatible changes" section of the release notes! github.com/rdkit/rdkit/...
Release 2026_03_1 (Q1 2026) Release · rdkit/rdkit
Release_2026.03.1 (Changes relative to Release_2025.09.1) Acknowledgements (Note: I'm no longer attempting to manually curate names. If you would like to see your contribution acknowledged with you...
github.com
Free registration for the 2026 #RDKit UGM (both in-person and online attendance) is now open: www.eventbrite.com/e/1985889262...
15th RDKit UGM 2026
2026 RDKit User Group Meeting (in person and online)
eventbrite.com
The new #RDKit blog post is a guest post looking at doing Butina clustering on the GPU using nvMolKit. greglandrum.github.io/rdkit-blog/p...
GPU-Accelerated Clustering with nvMolKit – RDKit blog
A guest post from NVIDIA
greglandrum.github.io
More terrible UI, this time from the @github.com app. I'm in landscape mode, vertical space is the limiter, but look at the incredible amount that is being wasted. Never mind that much of the info doesn't need to be displayed on that screen.
Oh, awesome. I can't wait to start finding hallucinated files in my OneDrive folders.
In this week's #RDKit blog post I assemble a new data set of compounds from patents. greglandrum.github.io/rdkit-blog/p...
Creating a patent data set – RDKit blog
Another collection of sets of related compounds
greglandrum.github.io
If I had any choice in the matter, I would stop using Word, PowerPoint, and Excel just because I am SO F*CKING TIRED of the constant "nudges" to use OneDrive.
I really enjoy my birthday, but this time I have had a particular Sammy Hagar song popping into my head for the last few days, and I'm very much ready to be done with it
The last #RDKit blog post of the year is a brief look back at 2025. greglandrum.github.io/rdkit-blog/p...
Wrapping up 2025 – RDKit blog
Completing the streak
greglandrum.github.io
I keep taking pictures like this in this gorge. Goodness do I love 🧗here.
Today's #RDKit blog post is a tutorial/explanation about which atoms are considered as candidates for tetrahedral chirality. greglandrum.github.io/rdkit-blog/p...
About tetrahedral chirality in the RDKit – RDKit blog
Answering a frequently asked question
greglandrum.github.io
Wow, the LinkedIn year in review was even dumber than I thought it was going to be. Maybe not everyone needs to do this.
@polarishub.io I'm starting a new project and wanted to use the Fang (a.k.a. Biogen) data set. Since I'm starting from scratch, this tim I figured I'd try grabbing it from polaris instead of doing my own curation. 1/n
This week's #RDKit blog post, like last week's, looks at creating your own synthon spaces. Last week was BRICS, this week we use some combichem reactions: greglandrum.github.io/rdkit-blog/p...
Building synthon spaces with combinatorial reactions – RDKit blog
A frequently requested tutorial part 2
greglandrum.github.io
This week's #RDKit blog post is an attempt to figure out thresholds for meaningful "similarity" with 3D methods. I think there may be a bit of tweaking to do here, but it's a start. greglandrum.github.io/rdkit-blog/p...
Thresholds for “random” with 3D similarity methods – RDKit blog
What is noise in 3D?
greglandrum.github.io
This week's #RDKit blog post is the third in a series using the LOBSTER data set. This time I use the data to compare 3D alignment methods. greglandrum.github.io/rdkit-blog/p...
Working with the LOBSTER Data set III – RDKit blog
Comparing 3D alignment methods.
greglandrum.github.io
Am I the only one getting a cloudflare error when they try to access ACS publications?