Riniker lab @ ETHZ

@rinikerlab.bsky.social

Riniker research group, ETH Zurich

Our new perspective in @jacs.acspublications.org makes some recommendations for building ML models for reaction outcomes. We didn't want to go too far out on a limb and call them "best practices" ;-) pubs.acs.org/doi/10.1021/...

Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes

Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.

pubs.acs.org

Our most recent paper just appeared in JCIM. The title of this one pretty much tells the story: when you assemble a data set by combining data from different literature assays, there is a very good chance that the resulting data contains a lot of noise. pubs.acs.org/doi/10.1021/...

Combining IC50 or Ki Values from Different Sources Is a Source of Significant Noise

As part of the ongoing quest to find or construct large data sets for use in validating new machine learning (ML) approaches for bioactivity prediction, it has become distressingly common for research...

pubs.acs.org