生成语法
计算机科学
人工智能
理论(学习稳定性)
机器学习
功能(生物学)
血凝素(流感)
生成模型
无监督学习
训练集
组分(热力学)
排名(信息检索)
蛋白质功能
实验数据
基础(线性代数)
偏爱
合成数据
蛋白质-蛋白质相互作用
作者
Talal Widatalla,Ashir A. Borah,S. B. King,C. Driscoll,Rafael Rafailov,Brian Hie
标识
DOI:10.1038/s41592-026-03137-3
摘要
Biological generative models can predict biological functions without task-specific training data but often under-perform specialized models. This is due to a fundamental 'alignment gap', where the rules learned during unsupervised training are not related to the function of interest. Here we demonstrate how to provide task-specific information without losing the general knowledge learned during pretraining by using direct preference optimization to align a structure-conditioned protein language model to preferentially generate stable protein sequences. Our aligned model, ProteinDPO, achieves stability prediction competitive to task-specific models and consistently outperforms unsupervised and fine-tuned versions of the model. Notably, ProteinDPO generalizes beyond its training data to enable stabilization and improved binding affinity prediction of large multichain protein complexes. When applied to stabilization of the hemagglutinin trimer, a primary component of influenza vaccines, ~80% of designs achieve increased or similar stability compared with the native hemagglutinin and up to 32 °C improvements from recently emerged mammalian strains. Our results demonstrate how to augment generative models with biophysical information and, more broadly, provide a general framework for the alignment of biological foundation models.
科研通智能强力驱动
Strongly Powered by AbleSci AI