Addressing dereplication crisis: Taxonomy-free reduction of massive genome collections using embeddings of protein content

Addressing dereplication crisis: Taxonomy-free reduction of massive genome collections using embeddings of protein content
复制标题

解决重复删除危机:使用蛋白质内容嵌入对大量基因组集合进行无分类减少

DOI:
10.1101/855262
复制
发表时间:
--
期刊:
bioRxiv
影响因子:
--
通讯作者:
​ C Brandt
​ C Brandt
中科院分区:
--
文献类型:
--
作者:
A Viehweger;M Hölzer;​ C Brandt

文献摘要

参考文献

相似文献

许多最近的微生物基因组收集整理了数十万个基因组。这本书使许多基因组分析变得复杂,例如分类单元分配,因为相关的计算负担是巨大的。然而,每个物种的代表数量都高度偏向于人类病原体和模式生物。因此,许多基因组包含的附加信息很少,可以被移除。我们创建了一种经济的去复制方法,可以仅基于基因组序列来减少大量的基因组集合,而不需要手动管理或分类信息。我们最近创建了细菌和古菌的基因组表示,称为“Nanotext”。这种方法将每个基因组嵌入一个低维的数字矢量中。扩展Nanotext,我们提出的算法Thinspace使用这些载体对相似基因组进行分组和去复制,我们将基因组分类数据库(GTDB)从大约15万个基因组去复制到不到2.2万个基因组。与NCBI RefSeq相比,由此产生的集合将元基因组数据集中的分类读取百分比提高了5倍,并且性能相当于更大的GTDB子集和手动管理的GTDB子集。使用Thinspace,可以在常规硬件上解除大规模基因组集合的复制,而不会影响下游结果。它是在BSD-3许可下发布的(githorb.com/phiweger/Thinspace)。
Many recent microbial genome collections curate hundreds of thousands of genomes. This volume complicates many genomic analyses such as taxon assignment because the associated computational burden is substantial. However, the number of representatives of each species is highly skewed towards human pathogens and model organisms. Thus many genomes contain little additional information and could be removed. We created a frugal dereplication method that can reduce massive genome collections based on genome sequence alone, without the need for manual curation nor taxonomic information.We recently created a genome representation for bacteria and archaea called “nanotext”. This method embeds each genome in a low-dimensional vector of numbers. Extending nanotext, our proposed algorithm called “thinspace” uses these vectors to group and dereplicate similar genomes.We dereplicated the Genome Taxonomy Database (GTDB) from about 150 thousand genomes to less than 22 thousand. The resulting collection increases the percent of classified reads in a metagenomic dataset by a factor of 5 compared to NCBI RefSeq and performs equal to both a larger as well as a manually curated GTDB subset.With thinspace, massive genome collections can be dereplicated on regular hardware, without affecting downstream results. It is released under a BSD-3 license (github.com/phiweger/thinspace).
对90K核基因组的高通量ANI分析揭示了清晰的物种边界。
DOI: 10.1038/s41467-018-07641-9
发表时间: 2018-11-30
影响因子: 16.6
作者:
Jain C;Rodriguez-R LM;Phillippy AM;Konstantinidis KT;Aluru S
通讯作者: Aluru S