BiG-SLiCE: A highly scalable tool maps the diversity of 1.2 million biosynthetic gene clusters.

BiG-SLiCE: A highly scalable tool maps the diversity of 1.2 million biosynthetic gene clusters.
复制标题

BiG-SLiCE:一种高度可扩展的工具,可绘制120万个生物合成基因簇的多样性。

DOI:
10.1093/gigascience/giaa154
复制
发表时间:
2021-01-13
期刊:
影响因子:
9.2
通讯作者:
Medema MH
Medema MH
中科院分区:
生物学2区
文献类型:
--
作者:
Kautsar SA;van der Hooft JJJ;de Ridder D;Medema MH

文献摘要

参考文献

被引文献

相似文献

生物合成基因簇(BGCs)的基因组挖掘已成为天然产物发现的重要组成部分。目前公开的20万个微生物基因组包含着丰富的新化学信息。导航这种巨大的基因组多样性的一种方法是通过同源bgc的比较分析,这允许识别跨物种模式,可以与代谢物或生物活动的存在相匹配。然而,目前的工具受到瓶颈的阻碍,瓶颈是用于将这些bgc分组为基因簇家族(gcf)的昂贵的基于网络的方法所造成的。在这里,我们介绍BiG-SLiCE,这是一个用于聚集大量bgc的工具。通过在欧几里得空间中表示它们,BiG-SLiCE可以以非成对、近线性的方式将bgc分组为gcf。我们使用BiG-SLiCE在一个典型的36核CPU服务器上分析了10天内从209,206个公开可用的微生物基因组和宏基因组组装基因组中收集的1,225,071个bgc。我们通过重建跨分类学的次级代谢多样性的全球地图来确定未知的生物合成潜力,从而证明了这种分析的实用性。BiG-SLiCE还提供了一个“查询模式”,可以有效地将新测序的bgc放入先前计算的gcf中,再加上一个强大的输出可视化引擎,方便用户友好的数据探索。BiG-SLiCE为加速天然产物的发现开辟了新的可能性,并为构建全球可搜索的bgc互联网络迈出了第一步。随着更多的基因组从未被充分研究的分类群中测序,更多的信息可以被挖掘出来,以突出它们潜在的新化学成分。BiG-SLiCE可以通过https://github.com/medema-group/bigslice获得。
Genome mining for biosynthetic gene clusters (BGCs) has become an integral part of natural product discovery. The >200,000 microbial genomes now publicly available hold information on abundant novel chemistry. One way to navigate this vast genomic diversity is through comparative analysis of homologous BGCs, which allows identification of cross-species patterns that can be matched to the presence of metabolites or biological activities. However, current tools are hindered by a bottleneck caused by the expensive network-based approach used to group these BGCs into gene cluster families (GCFs). Here, we introduce BiG-SLiCE, a tool designed to cluster massive numbers of BGCs. By representing them in Euclidean space, BiG-SLiCE can group BGCs into GCFs in a non-pairwise, near-linear fashion. We used BiG-SLiCE to analyze 1,225,071 BGCs collected from 209,206 publicly available microbial genomes and metagenome-assembled genomes within 10 days on a typical 36-core CPU server. We demonstrate the utility of such analyses by reconstructing a global map of secondary metabolic diversity across taxonomy to identify uncharted biosynthetic potential. BiG-SLiCE also provides a “query mode” that can efficiently place newly sequenced BGCs into previously computed GCFs, plus a powerful output visualization engine that facilitates user-friendly data exploration. BiG-SLiCE opens up new possibilities to accelerate natural product discovery and offers a first step towards constructing a global and searchable interconnected network of BGCs. As more genomes are sequenced from understudied taxa, more information can be mined to highlight their potentially novel chemistry. BiG-SLiCE is available via https://github.com/medema-group/bigslice.
DOI: 10.1073/pnas.1714381115
发表时间: 2017-12-26
影响因子: 11.1
作者:
Amos, Gregory C. A.;Awakawa, Takayoshi;Jensen, Paul R.
通讯作者: Jensen, Paul R.
DOI: 10.1038/s42003-019-0333-6
发表时间: 2019-02-28
影响因子: 5.9
作者:
Del Carratore,Francesco;Zych,Konrad;Breitling,Rainer
通讯作者: Breitling,Rainer
DOI: 10.1093/gbe/evw125
发表时间: 2016-07-02
影响因子: 3.3
作者:
Cruz-Morales P;Kopp JF;Martínez-Guerrero C;Yáñez-Guerra LA;Selem-Mojica N;Ramos-Aboites H;Feldmann J;Barona-Gómez F
通讯作者: Barona-Gómez F
DOI: 10.1186/s12859-017-1519-x
发表时间: 2017-02-13
期刊: BMC bioinformatics
影响因子: 3
作者:
Alborzi SZ;Devignes MD;Ritchie DW
通讯作者: Ritchie DW
DOI: 10.1186/1471-2148-10-26
发表时间: 2010-01-26
影响因子: 3.4
作者:
Bushley KE;Turgeon BG
通讯作者: Turgeon BG