Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions.
Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions.
复制标题
DOI:
10.1371/journal.pcbi.1008920
复制
发表时间:
2021-05
影响因子:
4.3
通讯作者:
Rogers S
中科院分区:
文献类型:
--
作者:
Hjörleifsson Eldjárn G;Ramsay A;van der Hooft JJJ;Duncan KR;Soldatou S;Rousu J;Daly R;Wandy J;Rogers S
Specialised metabolites from microbial sources are well-known for their wide range of biomedical applications, particularly as antibiotics. When mining paired genomic and metabolomic data sets for novel specialised metabolites, establishing links between Biosynthetic Gene Clusters (BGCs) and metabolites represents a promising way of finding such novel chemistry. However, due to the lack of detailed biosynthetic knowledge for the majority of predicted BGCs, and the large number of possible combinations, this is not a simple task. This problem is becoming ever more pressing with the increased availability of paired omics data sets. Current tools are not effective at identifying valid links automatically, and manual verification is a considerable bottleneck in natural product research. We demonstrate that using multiple link-scoring functions together makes it easier to prioritise true links relative to others. Based on standardising a commonly used score, we introduce a new, more effective score, and introduce a novel score using an Input-Output Kernel Regression approach. Finally, we present NPLinker, a software framework to link genomic and metabolomic data. Results are verified using publicly available data sets that include validated links. In this article, we introduce NPLinker, a software framework to link genomic and metabolomic data, to link microbial secondary metabolites to their producing genomic regions. Two of the major approaches for such linking are analysis of the correlation between sets of strains, and analysis of predicted features of the molecules. While these methods are usually used separately, we demonstrate that they are in fact complementary, and show a way to combine them to improve their performance. We begin by demonstrating a weakness in the most common method of strain correlation analysis, and suggest an improvement. We then introduce a new feature-based analysis method which, unlike most such methods, does not directly depend on the natural product compound class. Finally, we demonstrate that the two are complementary and proceed to combine them into a single scoring function for genomic and metabolomic links, which shows improved performance over either of the individual approaches. Verification is done using curated databases of genomic and metabolomic data, as well as public data sets of microbial data including validated links.
登录
查看更多内容
影响因子:
8.6
作者:
Heller SR;McNaught A;Pletnev I;Stein S;Tchekhovskoi D
通讯作者:
Tchekhovskoi D
影响因子:
9.2
作者:
Kautsar SA;van der Hooft JJJ;de Ridder D;Medema MH
通讯作者:
Medema MH
影响因子:
4
作者:
Mohimani, Hosein;Kersten, Roland D.;Liu, Wei-Ting;Wang, Mingxun;Purvine, Samuel O.;Wu, Si;Brewer, Heather M.;Pasa-Tolic, Ljiljana;Bandeira, Nuno;Moore, Bradley S.;Pevzner, Pavel A.;Dorrestein, Pieter C.
通讯作者:
Dorrestein, Pieter C.
影响因子:
--
作者:
Duncan KR;Crüsemann M;Lechner A;Sarkar A;Li J;Ziemert N;Wang M;Bandeira N;Moore BS;Dorrestein PC;Jensen PR
通讯作者:
Jensen PR
影响因子:
9.3
作者:
Cao, Liu;Gurevich, Alexey;Mohimani, Hosein
通讯作者:
Mohimani, Hosein