反应性(心理学)
外推法
公制(单位)
化学
共单体
计算机科学
统计
回归
回归分析
背景(考古学)
多元统计
数学
线性回归
计量经济学
共聚物
灵敏度(控制系统)
集合(抽象数据类型)
航程(航空)
人工智能
热力学
作者
Caroline M. Coxwell,Dylan M. Anstine,Olexandr Isayev,Frank A. Leibfarth
标识
DOI:10.1021/acs.macromol.6c01785
摘要
Abstract Reactivity ratios are a key metric for understanding copolymer microstructure, yet they are challenging to predict a priori. Data-driven methods have recently been employed for the prediction of free radical copolymerization reactivity ratios, but model extrapolation remains modest with respect to accuracy. This has elicited discussions on variability within published reactivity ratio datasets and a call for standardization. Here, we seek to identify the source of accuracy limitation in data-driven reactivity ratio prediction through the addition of new model features, evaluation of the model strategy, and estimate of the quality of literature reactivity ratio data. Systematic studies using a literature-mined dataset of >450 reactivity ratios demonstrated that the inclusion of more relevant transition state features in multivariate linear regression models and the use of more complex machine learning models for reactivity ratios led to marginal improvements in model accuracy, evaluated through mean absolute error. This motivated the evaluation of the accuracy of literature-extracted reactivity ratios through the analysis of a representative library of 100 reactivity ratio values from 47 distinct studies. Independent measurements of the same comonomer pair disagree by an average standard deviation of 0.26 kcal/mol in ΔΔG‡, an estimate of the irreducible noise in the training labels. The error of every model evaluated here, and most literature models, is comparable to this dispersion. Therefore, the study concludes that the accuracy floor in data-driven methods for reactivity ratio prediction is set by the data rather than by the descriptors or architecture and improving predictive models for reactivity ratios will require a large, standardized dataset of reactivity ratios measured using modern synthesis, characterization, and statistical analysis. More generally, the concept presented here of estimating the measurement–noise floor of a literature-mined dataset provides an approach to determine whether a prediction task is limited by the model or by the data, and we propose that it can be generalized beyond reactivity ratios to polymer properties that are compiled from heterogeneous literature.
科研通智能强力驱动
Strongly Powered by AbleSci AI