生物信息学
计算机科学
药物发现
冗余(工程)
机器学习
人工智能
计算模型
仿形(计算机编程)
训练集
任务(项目管理)
计算生物学
钥匙(锁)
一般化
Web服务器
数据整理
软件部署
集合(抽象数据类型)
数据建模
数据挖掘
虚拟筛选
数据科学
基于生理学的药代动力学模型
生物学数据
预测值
服务器
数量结构-活动关系
作者
Pierre Llompart,Claire Minoletti,Gilles Marcou,Alexandre Varnek
标识
DOI:10.1021/acs.jmedchem.6c00049
摘要
Multitask learning is a promising strategy in computational drug discovery, potentially improving predictive performance and generalization over traditional single-task models. MTL has shown particular value in absorption, distribution, metabolism, elimination, and toxicity (ADMET) and potency predictions, which are key for drug design. Yet, many existing Web servers rely on the same uncurated, decade-old data sets, creating an illusion of diversity. This work critically reviews open-source ADMET Web services, revealing extensive data redundancy and limited curation across the field. We introduce OneADMET, a meticulously curated data set of 738,161 compounds with 1,119,719 measurements spanning 44 ADMET end points and 1 489 biological activities. We report a unified ChemProp-based MTL model capable of handling hundreds of continuous tasks simultaneously, which has practical advantages for model deployment and maintenance. Additionally, we observed that these MTL models match or surpass single-task models in predictive accuracy. This study highlights the utility of large-scale MTL for pharmacokinetics profiling and contributes practical tools and data sets for the community.
科研通智能强力驱动
Strongly Powered by AbleSci AI