序列(生物学)
公制(单位)
人工智能
计算机科学
余弦相似度
瓶颈
相似性(几何)
语言模型
随机森林
计算生物学
结合位点
编码器
比例(比率)
蛋白质测序
机器学习
蛋白质-蛋白质相互作用
模式识别(心理学)
潜变量
交叉口(航空)
自然语言
自然语言处理
生物素化
肽序列
数学
序列比对
序列标记
化学
空格(标点符号)
生物系统
算法
作者
Harshit Singh,RAJEEV KUMAR SINGH,Satya Pratik Srivastava,Suryavedha Pradhan,Rohan Gorantla
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-04-01
被引量:1
标识
DOI:10.64898/2026.03.30.715237
摘要
Abstract Protein–protein interactions underpin virtually every aspect of cellular life, and the precise quantification of their binding affinity is fundamental to understanding immune recognition, disease mechanisms, and the rational design of therapeutic antibodies. Yet predicting binding affinity at scale remains an unsolved challenge: reliable experimental assays are low-throughput and expensive, while computational methods that depend on three-dimensional complex struc-tures cannot be applied to the vast majority of clinically relevant targets where structural data are absent. Here we present BALM-PPI, a framework that predicts protein–protein binding affinity from amino acid sequence alone. Both proteins are encoded by a protein language model trained on evolutionary sequence data and projected into a shared representational space, where their distance directly reflects binding strength. Fine-tuning this protein language model requires updating fewer than 1% of its parameters, and we show that this targeted adaptation steers the model toward interface-relevant sequence signals rather than spurious background correlations. On a curated benchmark of over 12,000 protein complexes, BALM-PPI matches or exceeds the accuracy of structure-based methods and retains predictive power for proteins with less than 30% sequence identity to the training set. Using only a subset of project-specific assay data, BALM-PPI outperforms a recent method trained on three times the data, suggesting that the model has already encoded the underlying interaction signals and requires only minimal supervision to specialise to a new target. BALM-PPI further provides residue-level attribution maps that pinpoint the amino acid positions driving each affinity prediction, consistently re-covering experimentally validated interaction hotspots across enzyme-inhibitor, signalling, and antibody-antigen systems without any structural input during training. This allows predictions to be cross-validated against structural and mutagenesis evidence, providing a mechanistic basis for candidate shortlisting ahead of experimental follow-up. BALM-PPI is freely accessible via an interactive web server.
科研通智能强力驱动
Strongly Powered by AbleSci AI