人工智能
深度学习
机器学习
标杆管理
计算机科学
基础(证据)
水准点(测量)
航程(航空)
随机森林
桥(图论)
分子描述符
深层神经网络
人工神经网络
不确定度量化
监督学习
代表(政治)
大数据
扩展(谓词逻辑)
强化学习
作者
Jackson Burns,Akshat Shirish Zalte,Charlles R. A. Abreu,Jochen Sieg,Christian Feldmann,Miriam Mathea,William H. Green
标识
DOI:10.1021/acs.jcim.6c01546
摘要
Abstract Fast and accurate data-driven prediction of molecular properties is pivotal to scientific advancements across myriad chemical domains. Deep learning methods have recently garnered much attention, despite their inability to outperform classical machine learning methods when tested on practical, real-world benchmarks with limited training data. This study seeks to bridge this gap by introducing a new avenue for foundation model pretraining. We propose pretraining on low-noise, calculable molecular descriptors via supervised learning to obtain rich, highly transferable molecular representations. We demonstrate this strategy with CheMeleon, a O(10M) parameter foundation model that enables directed message-passing neural networks to finally exceed the performance of classical methods in the low-data regime. We evaluate on 58 benchmark data sets spanning a range of properties relevant to small-molecule drug discovery, sourced from the industry-led Polaris benchmarking initiative. Rigorous statistical comparisons show that CheMeleon outperforms classical baselines like Random Forest on molecular fingerprints and descriptors, as well as existing foundation models. We open-source the CheMeleon model and the pretraining framework to encourage adoption and extension of this pretraining strategy across chemical sciences.
科研通智能强力驱动
Strongly Powered by AbleSci AI