反向
偏爱
多样性(政治)
数学优化
折叠(DSP实现)
计算机科学
数学
工程类
统计
政治学
结构工程
几何学
法学
作者
Ryan S. Park,Darren J. Hsu,C. Roland,Maria Korshunova,Chen Tessler,Shie Mannor,Olivia Viessmann,Bruno Trentini
标识
DOI:10.48550/arxiv.2410.19471
摘要
Inverse folding models play an important role in structure-based design by predicting amino acid sequences that fold into desired reference structures. Models like ProteinMPNN, a message-passing encoder-decoder model, are trained to reliably produce new sequences from a reference structure. However, when applied to peptides, these models are prone to generating repetitive sequences that do not fold into the reference structure. To address this, we fine-tune ProteinMPNN to produce diverse and structurally consistent peptide sequences via Direct Preference Optimization (DPO). We derive two enhancements to DPO: online diversity regularization and domain-specific priors. Additionally, we develop a new understanding on improving diversity in decoder models. When conditioned on OpenFold generated structures, our fine-tuned models achieve state-of-the-art structural similarity scores, improving base ProteinMPNN by at least 8%. Compared to standard DPO, our regularized method achieves up to 20% higher sequence diversity with no loss in structural similarity score.
科研通智能强力驱动
Strongly Powered by AbleSci AI