杠杆(统计)
数据科学
计算机科学
蛋白质功能
分类学(生物学)
功能(生物学)
光学(聚焦)
计算生物学
蛋白质测序
钥匙(锁)
人类蛋白质
知识管理
作者
Yang Xiao,Wanjia Zhao,Junkai Zhang,Yiqiao Jin,Han Zhang,Zhicheng Ren,Renliang Sun,Haixin Wang,Guangquan Wan,Lu Pan,Xiao‐Jun Luo,Yu Zhang,James Zou,Yizhou Sun,Wei Wang
标识
DOI:10.48550/arxiv.2502.17504
摘要
Protein-specific large language models (Protein LLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of Protein LLMs, covering their architectures, training datasets, evaluation metrics, and diverse applications. Through a systematic analysis of over 100 articles, we propose a structured taxonomy of state-of-the-art Protein LLMs, analyze how they leverage large-scale protein sequence data for improved accuracy, and explore their potential in advancing protein engineering and biomedical research. Additionally, we discuss key challenges and future directions, positioning Protein LLMs as essential tools for scientific discovery in protein science. Resources are maintained at https://github.com/Yijia-Xiao/Protein-LLM-Survey.
科研通智能强力驱动
Strongly Powered by AbleSci AI