工作流程
可扩展性
计算机科学
集合(抽象数据类型)
数据共享
数据集
数据建模
数据挖掘
试验装置
体积热力学
联合学习
数据模型(GIS)
数据库
概念证明
机器学习
班级(哲学)
合成数据
分布式计算
数据驱动
数据科学
预测建模
实验数据
分布式数据库
知识共享
大数据
人工智能
数据存取
试验数据
作者
Rajarshi Guha,Wenyi Wang,Edward Price,Barun Bhhatarai,Majdi Hassan,Anthony DiFranzo,Christopher Keefer,Nathaniel Woody,Fabio Broccatelli,Susanne Winiwarter,Ling He,Darren V. S. Green,Christopher D. Edwards,Harutoshi Kato,Yohei Kosugi,Prashant Desai
标识
DOI:10.1021/acs.jmedchem.5c03681
摘要
Machine learning models for ADMET prediction benefit from large, diverse data sets, yet such data are typically siloed across organizations. Federated learning (FL) enables collaborative modeling while preserving data privacy. Here, we investigate a student-teacher model (STM) framework in which organizations train internal models on proprietary data and share predictions on a public data set to generate pseudolabels for a centralized student model. As a proof of concept, 11 pharmaceutical companies contributed predictions for rat steady-state volume of distribution, yielding a pseudolabeled data set of ∼133,000 compounds. The resulting student model achieved performance comparable to individual teacher models on an external test set (RMSE ≈ 0.51 vs 0.47-0.61). Compared with FL approaches such as MELLODY and Effiris, STM offers a simpler workflow that avoids direct data sharing or iterative collaboration, providing a practical and scalable framework for secure cross-company model development.
科研通智能强力驱动
Strongly Powered by AbleSci AI