计算机科学
可扩展性
限制
代表(政治)
钥匙(锁)
选择(遗传算法)
编码器
编码(内存)
数据挖掘
机器学习
理论(学习稳定性)
人工智能
组分(热力学)
算法
特征选择
选型
期限(时间)
理论计算机科学
外部数据表示
数据建模
校准
一致性(知识库)
布线(电子设计自动化)
作者
Huiming Bao,Shouliang Dong
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-07-28
标识
DOI:10.64898/2026.07.24.740495
摘要
Abstract Protein-ligand affinity (PLA) prediction is central to AI-driven drug discovery, but precise interaction-based methods require costly conformation preparation and data encoding, limiting their throughput. To reconcile accuracy with efficiency, we first investigate whether pre-trained molecular representation models can replace complex encoders. A unified and diverse assessment of sequence-, graph-, and image-based representations reveals both strong overall performance and family-wise variability, delivering the first practical guidance for encoder selection in PLA tasks. Next, to achieve high computational efficiency without sacrificing expressiveness, we adopt the mixture-of-experts (MoE) strategy from large language models. Systematic ablation studies uncover key design principles for deploying MoE in molecular prediction. The resulting model, HydrAffinity, is an interaction-free, dynamic sparse model that uses pre-trained encoders and MoE for parameter-efficient learning. It outperforms all interaction-free methods and matches state-of-the-art interaction-based methods on CASF-2016. Routing analysis confirms that MoE develops distinct, family-specific activation patterns, providing interpretable evidence of dynamic parameterization across protein classes. Zero-shot tests on DUDE-Z and LIT-PCBA further show strong EF 5 % performance, making HydrAffinity a practical, scalable solution acting as an effective early-stage pre-filter. Graphic abstract
科研通智能强力驱动
Strongly Powered by AbleSci AI