相似性(几何)
计算机科学
化学信息学
人工智能
机器学习
任务(项目管理)
稀释
生物系统
训练集
人工神经网络
数据挖掘
试验装置
集合(抽象数据类型)
功能(生物学)
图形
试验数据
数量结构-活动关系
财产(哲学)
模式识别(心理学)
适用范围
数据集
相关系数
选型
活度系数
生化工程
选择(遗传算法)
实验数据
适应(眼睛)
主成分分析
数学
作者
Karol Baran,Adam Kloskowski
标识
DOI:10.1021/acs.jcim.6c00067
摘要
Collection of substantial training sets for structure–property modeling often poses a significant challenge, especially for niche sorption systems. This study provides computational experiments with meta-learning that presents a compelling solution to the data scarcity challenge in molecular property prediction. Our analysis utilized infinite dilution activity coefficients for several ionic liquid–solute systems, with coefficient prediction for systems with a particular solute constituting a task. The systems were modeled predominantly using model-agnostic meta-learning (MAML), supported by a study on Reptile and its modified variants. The obtained results provide promising insights into training MAML models by expanding the adaptation set size. Metrics such as R 2, RMSE, and MAE indicate comparable performance to graph neural networks, even when trained on only 64 or 128 data points. The versatility of the fine-tuned models suggests that, in certain cases, performance on a single task may be achieved at the expense of reduced model versatility (catastrophic forgetting). The similarity between the test and training tasks (as approximated by Tanimoto similarity of the solutes’ molecules) was identified as a factor affecting performance on the test task. Consequently, Task Similarity-Aware Reptile (TSA-Reptile) was proposed to target those dissimilar tasks. This novel method scales the loss function by similarity to the closest training task. It was shown to outperform MAML on out-of-distribution tasks. Beyond providing the comparative analysis of meta-learning and traditional deep learning, potential strengths of both MAML and TSA-Reptile are discussed.
科研通智能强力驱动
Strongly Powered by AbleSci AI