Our new perspective in @jacs.acspublications.org makes some recommendations for building ML models for reaction outcomes. We didn't want to go too far out on a limb and call them "best practices" ;-) pubs.acs.org/doi/10.1021/...
Yield Smarter, Not Harder: Good Practices for Machine Learning of Reaction Outcomes
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald–Hartwig (BH) amination, Suzuki–Miyaura (SM) coupling, and the silicon–amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
pubs.acs.org