残留物(化学)
化学
计算机科学
生物系统
计算生物学
人工智能
数据挖掘
蛋白质-蛋白质相互作用
序列(生物学)
算法
作者
Yun Zhou,Hanlei Qiu,Dong Liu,W M Wang
标识
DOI:10.1021/acs.jcim.6c01246
摘要
Accurate identification of DNA-binding proteins (DBPs) and their DNA-binding residue sites (DBSs) is essential for understanding gene regulatory processes. Despite recent progress achieved by protein language models, current methods still face fundamental limitations, including the quadratic computational burden of Transformer architectures, inadequate modeling of long-range dependencies, and reduced generalization on long or low-homology protein sequences. To address these challenges, we propose ESM2-BiMamba, a length-adaptive hybrid architecture for efficient and scalable protein-DNA interaction prediction. The model preserves the first 29 Transformer layers of the pretrained 33-layer ESM2 and replaces its top four layers with bidirectional Mamba state-space modules, enabling linear-time context propagation while maintaining rich sequence semantics. A sequence-length-adaptive dynamic chunking mechanism further reduces redundant computation and stabilizes long-range dependency modeling. To mitigate the distribution shift between pretraining and downstream tasks, a lightweight adapter is incorporated to enhance representation alignment. In addition, a dual-task prediction head comprising a protein-level DBP classifier and a residue-level DBS predictor allows the model to jointly capture global functional patterns and fine-grained binding-site signals. Extensive experiments on multiple standard benchmark data sets demonstrate that ESM2-BiMamba achieves superior performance in both DNA-binding protein identification and DNA-binding residue prediction, with notable advantages in processing long sequences and generalizing to low-homology targets.
科研通智能强力驱动
Strongly Powered by AbleSci AI