The impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters

The impact of contaminants on the accuracy of genome skimming and the effectiveness of exclusion read filters
复制标题

DOI:
10.1111/1755-0998.13135
复制
发表时间:
2020-02-04
影响因子:
7.7
通讯作者:
Mirarab, Siavash
Mirarab, Siavash
中科院分区:
生物学1区
文献类型:
--
作者:
Rachtman, Eleonora;Balaban, Metin;Mirarab, Siavash

文献摘要

被引文献

相似文献

检测从环境中获得的样本的身份的能力是分子生态学研究的基石。由于散弹枪测序的价格不断下降,基因组略读,即在低覆盖率下获取跨基因组的短读数,正在成为传统条形码的替代方案。通过在整个基因组中获得更多的数据,略读技术有望在保持成本可控的同时,提高样本识别的精度,而不是传统的条形码。虽然现在已经有了基于基因组框架的免组装样本鉴定方法,但人们对这些方法对来自目标物种以外的生物的DNA的反应知之甚少。在这篇文章中,我们证明了基于k-mer相似性计算的一对基因组撇子之间的距离的准确性会显著降低,如果这些撇子包括污染读数,即任何来自其他生物的读数。建立了污染影响的理论模型。然后,我们建议和评估污染问题的解决方案:根据可能的污染物(例如,所有微生物)的广泛数据库,查询读取基因组,并过滤出任何匹配的读取。我们在详细的分析中评估了这一战略在使用Kraken-II实施时的有效性。我们的结果表明,由于过滤的结果,准确度有了实质性的提高,但也指出了局限性,包括需要在污染物数据库中进行相对接近的匹配。
The ability to detect the identity of a sample obtained from its environment is a cornerstone of molecular ecological research. Thanks to the falling price of shotgun sequencing, genome skimming, the acquisition of short reads spread across the genome at low coverage, is emerging as an alternative to traditional barcoding. By obtaining far more data across the whole genome, skimming has the promise to increase the precision of sample identification beyond traditional barcoding while keeping the costs manageable. While methods for assembly-free sample identification based on genome skims are now available, little is known about how these methods react to the presence of DNA from organisms other than the target species. In this paper, we show that the accuracy of distances computed between a pair of genome skims based on k-mer similarity can degrade dramatically if the skims include contaminant reads; i.e., any reads originating from other organisms. We establish a theoretical model of the impact of contamination. We then suggest and evaluate a solution to the contamination problem: Query reads in a genome skim against an extensive database of possible contaminants (e.g., all microbial organisms) and filter out any read that matches. We evaluate the effectiveness of this strategy when implemented using Kraken-II, in detailed analyses. Our results show substantial improvements in accuracy as a result of filtering but also point to limitations, including a need for relatively close matches in the contaminant database.