瓶颈
计算机科学
排名(信息检索)
数据挖掘
代谢组学
任务(项目管理)
关联规则学习
基因组学
计算生物学
软件
机器学习
生物信息学
生物
基因组
基因
工程类
遗传学
程序设计语言
系统工程
嵌入式系统
作者
Grímur Hjörleifsson Eldjárn,Andrew Ramsay,Justin J. J. van der Hooft,Katherine Duncan,Sylvia Soldatou,Juho Rousu,Rónán Daly,Joe Wandy,Simon Rogers
标识
DOI:10.1371/journal.pcbi.1008920
摘要
Specialised metabolites from microbial sources are well-known for their wide range of biomedical applications, particularly as antibiotics. When mining paired genomic and metabolomic data sets for novel specialised metabolites, establishing links between Biosynthetic Gene Clusters (BGCs) and metabolites represents a promising way of finding such novel chemistry. However, due to the lack of detailed biosynthetic knowledge for the majority of predicted BGCs, and the large number of possible combinations, this is not a simple task. This problem is becoming ever more pressing with the increased availability of paired omics data sets. Current tools are not effective at identifying valid links automatically, and manual verification is a considerable bottleneck in natural product research. We demonstrate that using multiple link-scoring functions together makes it easier to prioritise true links relative to others. Based on standardising a commonly used score, we introduce a new, more effective score, and introduce a novel score using an Input-Output Kernel Regression approach. Finally, we present NPLinker, a software framework to link genomic and metabolomic data. Results are verified using publicly available data sets that include validated links.
科研通智能强力驱动
Strongly Powered by AbleSci AI