广告
机器学习
计算机科学
人工智能
排名(信息检索)
优先次序
多任务学习
人工神经网络
学习迁移
训练集
深度学习
数据挖掘
图形
药物发现
实验数据
公制(单位)
代表(政治)
钥匙(锁)
稀缺
作者
Long-Hung Dinh Pham,Minh-Tri Le,Khac‐Minh Thai
标识
DOI:10.1021/acs.jcim.5c02030
摘要
Absorption, distribution, metabolism, and excretion (ADME) properties are among the key factors in determining the success of lead discovery and optimization campaigns. Fast and accurate prediction of molecular ADME profiles is hence of particular interest as a prioritization tool before costly experimental assays. However, the severe scarcity of publicly available training data for ADME prediction has hindered the development of improved machine learning models. Recently, industry teams have taken the important step to release the predicted labels from their in-house trained models for public domain chemical structures. In this paper, leveraging these large and diverse surrogate data sets, we propose the adoption of transfer learning using a simple multitask graph neural network (GNN) for rich representation learning and focused fine-tuning on experimental data. In participation of the blinded ASAP-Polaris-OpenADMET antiviral ADME challenge 2025, the approach achieved competitive results, ranking fourth on aggregated mean absolute error (MAE) and tied second on aggregated Pearson R. Post-competition optimization further pushed the performance to surpass the third-place entry in MAE, without using any proprietary data or commercial featurization methods. We further explored a pretraining strategy integrating both experimental and predicted labels, showing improvements and a promising direction for pretraining on data from multiple sources. The study presents an example of new opportunities for making use of predicted labels for pretraining and applications to real-world tasks. The code and pretrained models are available on: https://github.com/LongHung-Pham/pADME.
科研通智能强力驱动
Strongly Powered by AbleSci AI