共价键
化学
半胱氨酸
标杆管理
残留物(化学)
鉴定(生物学)
计算生物学
组合化学
生物化学
药物发现
共价结合
药物靶点
非共价相互作用
纳米技术
靶蛋白
药物开发
化学生物学
血浆蛋白结合
小分子
蛋白质结构
结构-活动关系
作者
Yanlin Ren,Minjie Mou,朱怡淼,Ziqi Pan,Kuo Zhang,Yuntao Qian,Yang Zhang,Jinlong Li,Tingting Fu,Feng Zhu
标识
DOI:10.1021/acs.jmedchem.6c01911
摘要
Abstract Targeted covalent inhibition is an important strategy in modern drug discovery, with cysteine being the most common residue targeted for covalent ligands. Accurate identification of covalently ligandable cysteines is therefore essential, especially for traditionally “undruggable” targets. However, structure-based methods depend on available and reliable protein structures, while sequence-based methods remain scarce and require further improvement. Here, we present CCSite, a protein language model-based framework for discovering covalently ligandable cysteines from protein sequences. It uniquely integrates low-rank adaptation of ESM Cambrian (LoRA-ESMC) with a cysteine-centered encoder-decoder module to capture local microenvironment features and long-range contextual information. Benchmarking and independent evaluation showed that CCSite achieved competitive performance without requiring 3D structures. Moreover, a real-world application revealed that CCSite was capable of prospectively identifying experimentally validated covalent cysteines, and a large-scale screen of human pathogenic X-to-Cys mutations further identified over 2,000 neo-cysteines as covalently ligandable candidates for future covalent drug development.
科研通智能强力驱动
Strongly Powered by AbleSci AI