计算机科学
桥接(联网)
代表(政治)
可扩展性
人工智能
自然语言
虚拟筛选
序列(生物学)
一般化
判别式
人机交互
自然语言处理
可视化
空格(标点符号)
接口(物质)
建筑
用户界面
药物发现
化学空间
融合
理论计算机科学
计算模型
机器学习
中间语言
语言模型
外部数据表示
作者
Rui Gu,Yingxu Liu,Bingxing Zhu,Li Liang,Haichun Liu,Yanmin Zhang,Yadong Chen
标识
DOI:10.1021/acs.jcim.5c02499
摘要
Traditional molecular screening methods are often limited by high computational cost, long design cycles, and a strong reliance on high-quality 3D protein structures, which are not always available or reliable. To address these limitations, we propose CoDrug, an innovative multimodal fusion framework that integrates textual information with structural representations of proteins and compounds. CoDrug employs two complementary fusion strategies─text-protein sequence fusion, in which SciBERT encodes functional descriptions and ESM extracts sequence-level features, and text-compound structure fusion, in which ChemFormer encodes SMILES and SciBERT processes compound-related textual descriptions. Using contrastive learning, CoDrug aligns textual and structural embeddings in a shared latent space, enabling effective cross-modal representation learning. This architecture supports novel functionalities, including text-driven virtual screening and text-driven molecular optimization, enhancing representation expressiveness and generalization while delivering strong performance under zero-shot settings. Evaluations on diverse benchmarks demonstrate that CoDrug achieves competitive or superior results compared with state-of-the-art baselines, particularly when 3D structural data are incomplete or unavailable. The framework's natural language interface lowers the technical barrier for AI-assisted drug discovery, allowing chemists to efficiently navigate and optimize chemical space without specialized computational expertise. By bridging language-driven hypotheses and structure-guided molecular design, CoDrug offers a scalable and flexible paradigm for accelerating the early stages of drug discovery.
科研通智能强力驱动
Strongly Powered by AbleSci AI