概化理论
稳健性(进化)
计算机科学
机器学习
水准点(测量)
人工智能
人工神经网络
试验数据
特征工程
数据挖掘
特征向量
深度学习
统计
数学
生物化学
化学
大地测量学
基因
程序设计语言
地理
作者
Kangming Li,Brian DeCost,Kamal Kumar Choudhary,Michael Greenwood,Jason R. Hattrick-Simpers
标识
DOI:10.1038/s41524-023-01012-9
摘要
Abstract Recent advances in machine learning (ML) have led to substantial performance improvement in material database benchmarks, but an excellent benchmark score may not imply good generalization performance. Here we show that ML models trained on Materials Project 2018 can have severely degraded performance on new compounds in Materials Project 2021 due to the distribution shift. We discuss how to foresee the issue with a few simple tools. Firstly, the uniform manifold approximation and projection (UMAP) can be used to investigate the relation between the training and test data within the feature space. Secondly, the disagreement between multiple ML models on the test data can illuminate out-of-distribution samples. We demonstrate that the UMAP-guided and query by committee acquisition strategies can greatly improve prediction accuracy by adding only 1% of the test data. We believe this work provides valuable insights for building databases and models that enable better robustness and generalizability.
科研通智能强力驱动
Strongly Powered by AbleSci AI