To Dereplicate or Not To Dereplicate?

To Dereplicate or Not To Dereplicate?
复制标题

DOI:
10.1128/msphere.00971-19
复制
发表时间:
2020-05-01
期刊:
影响因子:
4.8
通讯作者:
Denef, Vincent J.
Denef, Vincent J.
中科院分区:
生物学2区
文献类型:
--
作者:
Evans, Jacob T.;Denef, Vincent J.

文献摘要

被引文献

相似文献

元基因组组装(MAG)扩展了我们对微生物多样性、进化和生态学的理解。人们对测序、组装、装箱和质量评估工具如何导致MAG不能反映自然界中的单个种群提出了关注。在这里,我们思考另一个问题,即如何处理由独立数据集组装而成的高度相似的MAG。获得一个物种的多个基因组代表是非常有价值的,因为它允许进行种群基因组分析;然而,当保留密切相关种群的基因组时,它使MAG质量评估和丰度推断复杂化。我们表明:(I)已发表的数据集包含很大一部分MAG共享&>99%的平均核苷酸同一性;(Ii)用于解决这种冗余的不同软件包和参数移除了非常不同数量的MAG;以及(Iii)密切相关基因组的移除会导致群体特异性辅助基因的丢失。最后,我们强调了一些方法,这些方法可以在不去复制的情况下跨样本系列推断特定菌株的动态。
Metagenome-assembled genomes (MAGs) expand our understanding of microbial diversity, evolution, and ecology. Concerns have been raised on how sequencing, assembly, binning, and quality assessment tools may result in MAGs that do not reflect single populations in nature. Here, we reflect on another issue, i.e., how to handle highly similar MAGs assembled from independent data sets. Obtaining multiple genomic representatives for a species is highly valuable, as it allows for population genomic analyses; however, when retaining genomes of closely related populations, it complicates MAG quality assessment and abundance inferences. We show that (i) published data sets contain a large fraction of MAGs sharing >99% average nucleotide identity, (ii) different software packages and parameters used to resolve this redundancy remove very different numbers of MAGs, and (iii) the removal of closely related genomes leads to losses of population-specific auxiliary genes. Finally, we highlight some approaches that can infer strain-specific dynamics across a sample series without dereplication.