可解释性
机器学习
人工智能
稳健性(进化)
水准点(测量)
一般化
计算机科学
预测建模
训练集
支持向量机
产量(工程)
Boosting(机器学习)
泛化误差
人工神经网络
数据挖掘
实验数据
基线(sea)
交叉验证
计算模型
极限(数学)
合成数据
深度学习
集成学习
大数据
作者
Idil Ismail,Gregory A. Landrum,Sereina Riniker
摘要
Reaction yield prediction is a longstanding challenge in synthetic chemistry, with broad implications for route planning, scalability, and high-throughput experimentation (HTE). While recent machine learning (ML) approaches have demonstrated promise in modeling reactivity, they often use complex descriptors or deep architectures that are computationally expensive and limit interpretability and scalability. Here, we assess how much information is stored in simpler descriptors and whether model accuracy is improved by increasing the complexity of the descriptors. Using classical ML models trained on descriptors with different complexity levels, we benchmark predictive performance on four publicly available HTE data sets covering three diverse reaction data sets: Buchwald-Hartwig (BH) amination, Suzuki-Miyaura (SM) coupling, and the silicon-amine protocol (SLAP). Our evaluation furthermore discusses (1) generalization via component-wise data splitting, (2) robustness through external validation across data sets, and (3) performance across asymmetric yield distributions characteristic of HTE data. Contrary to conventional expectations, we find that simpler models with interpretable features can achieve competitive performance under rigorous validation protocols. Based on our findings, we formulate good practices for future studies in this area. For example, comparison to low-cost baseline models should become a requirement for future ML studies for reaction-yield prediction.
科研通智能强力驱动
Strongly Powered by AbleSci AI