可解释性
片段(逻辑)
化学空间
计算机科学
药物发现
限制
机器学习
人工智能
回顾性分析
钥匙(锁)
化学信息学
数据挖掘
合成数据
数量结构-活动关系
计算生物学
计算模型
变量(数学)
组合化学
药物开发
生化工程
小分子
训练集
化学数据库
化学
生物系统
有效载荷(计算)
特征(语言学)
药物靶点
药品
分子描述符
作者
Xiang Zhang,Jia Liu,Bufan Xu,Z. Conrad Zhang,Zifu Huang,Kaixian Chen,Dingyan Wang,Xutong Li
标识
DOI:10.1021/acs.jcim.5c02450
摘要
AI-driven molecular generation encounters a "generation-synthesis gap": most computationally designed molecules cannot be synthesized in laboratories, limiting AI-assisted drug design (AIDD) applications. Current approaches to assess synthetic accessibility (SA) include computer-aided synthesis planning (CASP) tools that perform retrosynthetic searches and machine learning-based SA prediction models that provide rapid scoring. CASP tools are computationally expensive for high-throughput screening, while existing SA prediction models may lack chemical synthesis logic or exhibit variable performance across different chemical spaces. We developed SynFrag, an SA prediction model using fragment assembly autoregressive generation to learn stepwise molecular construction patterns. Self-supervised pretraining on millions of unlabeled molecules enables the learning of dynamic fragment assembly patterns beyond fragment occurrence statistics or reaction step annotations. This approach captures connectivity relationships relevant to synthesis difficulty cliffs, where minor structural changes substantially alter SA. Evaluation across public benchmarks, clinical drugs with intermediates, and AI-generated molecules shows consistent performance across diverse chemical spaces. The model produces subsecond predictions with attention mechanisms corresponding to key reactive sites. SynFrag provides computational efficiency suitable for large-scale screening while maintaining interpretability for detailed SA assessment in drug discovery workflows. Online platform: https://synfrag.simm.ac.cn. Code and data available: https://github.com/simmzx/SynFrag.
科研通智能强力驱动
Strongly Powered by AbleSci AI