吩嗪
生物信息学
计算生物学
机器学习
人工智能
酶
计算机科学
生物化学
化学空间
化学
生物
药物发现
训练集
机制(生物学)
基因
底物特异性
化学信息学
功能基因组学
基质(水族馆)
基因组
基因组学
组合化学
鉴定(生物学)
蛋白质测序
特征(语言学)
生物信息学
遗传学
深度学习
作者
Xiaoyu Shan,Inês B. Trindade,Nathaniel R. Glasser,Korbinian O. Thalhammer,Matthew Scurria,Ariane Mora,Stuart J. Conway,Dianne K. Newman
出处
期刊:
[Cold Spring Harbor Laboratory]
日期:2026-03-06
标识
DOI:10.64898/2026.03.05.709892
摘要
Abstract Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. Applying this framework led to the discovery of PTC (Phenazine-Thiol Conjugase), the first enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through non-enzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert “small data” typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery. Significance Machine learning excels when large, well-labeled datasets are available, yet many biologically important problems lack sufficient experimental data to support such approaches to discovery. This limitation is particularly acute for identifying enzymes acting on rare or understudied substrates. Here, we show that genomic organization can be leveraged as an additional source of biological information to address data sparsity. Starting with only 14 enzymes experimentally shown to modify phenazines, we developed a model identifying phenazine-interacting enzymes by integrating genome-informed data augmentation with protein machine learning. Guided by the model, we discovered the first enzyme known to catalyze thioconjugation modifications of phenazines, demonstrating a simple yet powerful strategy for extracting predictive insight from sparse biological knowledge.
科研通智能强力驱动
Strongly Powered by AbleSci AI