Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions.

Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions.
复制标题

DOI:
10.1371/journal.pcbi.1008920
复制
发表时间:
2021-05
影响因子:
4.3
通讯作者:
Rogers S
Rogers S
中科院分区:
生物学2区
文献类型:
--
作者:
Hjörleifsson Eldjárn G;Ramsay A;van der Hooft JJJ;Duncan KR;Soldatou S;Rousu J;Daly R;Wandy J;Rogers S

文献摘要

参考文献

被引文献

相似文献

来自微生物来源的专门代谢物以其广泛的生物医学应用而闻名,特别是作为抗生素。当挖掘成对的基因组和代谢组学数据集以寻找新的专门代谢物时,在生物合成基因簇(BGC)和代谢物之间建立联系代表了找到这种新化学的有希望的方法。然而,由于对大多数预测的BGC缺乏详细的生物合成知识,以及大量可能的组合,这不是一个简单的任务。随着配对组学数据集可用性的增加,这个问题变得越来越紧迫。目前的工具不能有效地自动识别有效的链接,人工验证是天然产物研究的一个相当大的瓶颈。我们证明,使用多个链接评分功能一起使用,可以更容易地优先考虑真正的链接相对于他人。基于标准化常用的分数,我们引入了一个新的,更有效的分数,并引入了一个新的分数使用输入输出核回归方法。最后,我们提出了NPLinker,一个软件框架,以连接基因组和代谢组学数据。使用包括经验证链接的公开数据集对结果进行验证。在这篇文章中,我们介绍了NPLinker,一个软件框架,连接基因组和代谢组学数据,微生物次级代谢产物的生产基因组区域。两个主要的方法,这样的链接是分析之间的相关性的菌株,和分析的预测功能的分子。虽然这些方法通常是单独使用,我们证明,他们实际上是互补的,并显示了一种方法来联合收割机,以提高其性能。我们开始通过展示应变相关分析的最常见的方法的弱点,并提出改进。然后,我们介绍了一种新的基于特征的分析方法,不像大多数这样的方法,不直接依赖于天然产物化合物类。最后,我们证明了这两个是互补的,并继续将它们联合收割机组合成一个单一的评分功能的基因组和代谢组学的链接,这表明改进的性能超过任何一个单独的方法。使用基因组和代谢组学数据的策展数据库以及微生物数据的公共数据集(包括经验证的链接)进行验证。
Specialised metabolites from microbial sources are well-known for their wide range of biomedical applications, particularly as antibiotics. When mining paired genomic and metabolomic data sets for novel specialised metabolites, establishing links between Biosynthetic Gene Clusters (BGCs) and metabolites represents a promising way of finding such novel chemistry. However, due to the lack of detailed biosynthetic knowledge for the majority of predicted BGCs, and the large number of possible combinations, this is not a simple task. This problem is becoming ever more pressing with the increased availability of paired omics data sets. Current tools are not effective at identifying valid links automatically, and manual verification is a considerable bottleneck in natural product research. We demonstrate that using multiple link-scoring functions together makes it easier to prioritise true links relative to others. Based on standardising a commonly used score, we introduce a new, more effective score, and introduce a novel score using an Input-Output Kernel Regression approach. Finally, we present NPLinker, a software framework to link genomic and metabolomic data. Results are verified using publicly available data sets that include validated links. In this article, we introduce NPLinker, a software framework to link genomic and metabolomic data, to link microbial secondary metabolites to their producing genomic regions. Two of the major approaches for such linking are analysis of the correlation between sets of strains, and analysis of predicted features of the molecules. While these methods are usually used separately, we demonstrate that they are in fact complementary, and show a way to combine them to improve their performance. We begin by demonstrating a weakness in the most common method of strain correlation analysis, and suggest an improvement. We then introduce a new feature-based analysis method which, unlike most such methods, does not directly depend on the natural product compound class. Finally, we demonstrate that the two are complementary and proceed to combine them into a single scoring function for genomic and metabolomic links, which shows improved performance over either of the individual approaches. Verification is done using curated databases of genomic and metabolomic data, as well as public data sets of microbial data including validated links.
DOI: 10.1186/s13321-015-0068-4
发表时间: 2015
影响因子: 8.6
作者:
Heller SR;McNaught A;Pletnev I;Stein S;Tchekhovskoi D
通讯作者: Tchekhovskoi D
DOI: 10.1093/gigascience/giaa154
发表时间: 2021-01-13
期刊: GigaScience
影响因子: 9.2
作者:
Kautsar SA;van der Hooft JJJ;de Ridder D;Medema MH
通讯作者: Medema MH
DOI: 10.1021/cb500199h
发表时间: 2014-07-18
影响因子: 4
作者:
Mohimani, Hosein;Kersten, Roland D.;Liu, Wei-Ting;Wang, Mingxun;Purvine, Samuel O.;Wu, Si;Brewer, Heather M.;Pasa-Tolic, Ljiljana;Bandeira, Nuno;Moore, Bradley S.;Pevzner, Pavel A.;Dorrestein, Pieter C.
通讯作者: Dorrestein, Pieter C.
DOI: 10.1016/j.chembiol.2015.03.010
发表时间: 2015-04-23
影响因子: --
作者:
Duncan KR;Crüsemann M;Lechner A;Sarkar A;Li J;Ziemert N;Wang M;Bandeira N;Moore BS;Dorrestein PC;Jensen PR
通讯作者: Jensen PR
DOI: 10.1016/j.cels.2019.09.004
发表时间: 2019-12-18
期刊: CELL SYSTEMS
影响因子: 9.3
作者:
Cao, Liu;Gurevich, Alexey;Mohimani, Hosein
通讯作者: Mohimani, Hosein