HAYSTAC: A Bayesian framework for robust and rapid species identification in high-throughput sequencing data.

HAYSTAC: A Bayesian framework for robust and rapid species identification in high-throughput sequencing data.
复制标题

HAYSTAC:一个贝叶斯框架,用于在高通量测序数据中进行稳健和快速的物种识别。

DOI:
10.1371/journal.pcbi.1010493
复制
发表时间:
2022-09
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

宏基因组样品中特定物种的鉴定对于几个关键应用至关重要,然而许多可用的工具需要大量的计算能力,并且往往容易出现假阳性鉴定。在这里,我们描述了高精度和可扩展的宏基因组数据分类分配(HAYSTAC),它可以估计特定分类单元存在于宏基因组中的概率。HAYSTAC提供了一个用户友好的工具来构建数据库,基于公开可用的基因组,用于竞争性读取映射。然后,它使用一种新的贝叶斯框架来推断每个物种识别的丰度和统计支持,并提供每次读取的物种分类。与其他方法不同,HAYSTAC专门设计用于有效处理古代和现代DNA数据以及不完整的参考数据库,从而可以运行高度准确的假设驱动分析(即,评估特定物种的存在),同时显著提高处理速度。我们使用模拟的Illumina文库测试了HAYSTAC的性能和准确性,包括有和没有古DNA损伤,并将结果与其他目前可用的方法(即,Kraken2/Bracken、KrakenUniq、MALT/HOPS和Sigma)。HAYSTAC在所有模拟中识别的假阳性比Kraken 2/Bracken,KrakenUniq和MALT都少,在模拟古代数据时比Sigma少。在数据库构建和样本分析过程中,它使用的内存比Kraken 2/Bracken,KrakenUniq以及MALT更少。最后,我们使用HAYSTAC在两个已发表的古代宏基因组数据集中搜索特定病原体,展示了它如何应用于经验数据集。HAYSTAC可从https://github.com/antonisdim/HAYSTAC获得。新兴的古宏基因组学领域(即,宏基因组学(从古代DNA)在病原体进化和古环境重建等领域的新发现中具有巨大的前景。然而,目前缺乏从降解和非降解DNA材料中的微生物群落中识别物种的计算方法。在这里,我们提出了“HAYSTAC”,一个用户友好的软件包,实现了一种新的概率模型,从降解和非降解的DNA材料获得的宏基因组数据中的物种识别。通过广泛的基准测试,我们表明,HAYSTAC可以用于准确地分析社区组成,以及直接假设检验的存在下,极低丰度类群,在复杂的宏基因组样本。在分析模拟和公开可用的数据集后,HAYSTAC在分类学分析期间始终产生最低数量的假阳性鉴定,当使用有限大小的数据库时产生稳健的结果,并且与其他专业方法相比,显示出病原体检测的灵敏度增加。HAYSTAC采用的新提出的概率模型和软件可以对降解/浅测序宏基因组样本中的稳健和快速病原体发现产生重大影响,同时优化计算资源的使用。
Identification of specific species in metagenomic samples is critical for several key applications, yet many tools available require large computational power and are often prone to false positive identifications. Here we describe High-AccuracY and Scalable Taxonomic Assignment of MetagenomiC data (HAYSTAC), which can estimate the probability that a specific taxon is present in a metagenome. HAYSTAC provides a user-friendly tool to construct databases, based on publicly available genomes, that are used for competitive read mapping. It then uses a novel Bayesian framework to infer the abundance and statistical support for each species identification and provide per-read species classification. Unlike other methods, HAYSTAC is specifically designed to efficiently handle both ancient and modern DNA data, as well as incomplete reference databases, making it possible to run highly accurate hypothesis-driven analyses (i.e., assessing the presence of a specific species) on variably sized reference databases while dramatically improving processing speeds. We tested the performance and accuracy of HAYSTAC using simulated Illumina libraries, both with and without ancient DNA damage, and compared the results to other currently available methods (i.e., Kraken2/Bracken, KrakenUniq, MALT/HOPS, and Sigma). HAYSTAC identified fewer false positives than both Kraken2/Bracken, KrakenUniq and MALT in all simulations, and fewer than Sigma in simulations of ancient data. It uses less memory than Kraken2/Bracken, KrakenUniq as well as MALT both during database construction and sample analysis. Lastly, we used HAYSTAC to search for specific pathogens in two published ancient metagenomic datasets, demonstrating how it can be applied to empirical datasets. HAYSTAC is available from https://github.com/antonisdim/HAYSTAC. The emerging field of paleo-metagenomics (i.e., metagenomics from ancient DNA) holds great promise for novel discoveries in fields as diverse as pathogen evolution and paleoenvironmental reconstruction. However, there is presently a lack of computational methods for species identification from microbial communities in both degraded and nondegraded DNA material. Here, we present “HAYSTAC”, a user-friendly software package that implements a novel probabilistic model for species identification in metagenomic data obtained from both degraded and non-degraded DNA material. Through extensive benchmarking, we show that HAYSTAC can be used for accurately profiling the community composition, as well as for direct hypothesis testing for the presence of extremely low-abundance taxa, in complex metagenomic samples. After analysing simulated and publicly available datasets, HAYSTAC consistently produced the lowest number of false positive identifications during taxonomic profiling, produced robust results when databases of restricted size were used, and showed increased sensitivity for pathogen detection compared to other specialist methods. The newly proposed probabilistic model and software employed by HAYSTAC can have a substantial impact on the robust and rapid pathogen discovery in degraded/shallow sequenced metagenomic samples while optimising the use of computational resources.
DOI: 10.1038/nature10549
发表时间: 2011-10-12
期刊: NATURE
影响因子: 64.8
作者:
Bos, Kirsten I.;Schuenemann, Verena J.;Golding, G. Brian;Burbano, Hernan A.;Waglechner, Nicholas;Coombes, Brian K.;McPhee, Joseph B.;DeWitte, Sharon N.;Meyer, Matthias;Schmedes, Sarah;Wood, James;Earn, David J. D.;Herring, D. Ann;Bauer, Peter;Poinar, Hendrik N.;Krause, Johannes
通讯作者: Krause, Johannes
DOI: 10.1038/nmeth.1923
发表时间: 2012-03-04
期刊: NATURE METHODS
影响因子: 48
作者:
Langmead, Ben;Salzberg, Steven L.
通讯作者: Salzberg, Steven L.
DOI: 10.1186/s12864-020-07229-y
发表时间: 2020-11-30
期刊: BMC genomics
影响因子: 4.4
作者:
Feuerborn TR;Palkopoulou E;van der Valk T;von Seth J;Munters AR;Pečnerová P;Dehasque M;Ureña I;Ersmark E;Lagerholm VK;Krzewińska M;Rodríguez-Varela R;Götherström A;Dalén L;Díez-Del-Molino D
通讯作者: Díez-Del-Molino D
DOI: 10.1186/s40168-018-0605-2
发表时间: 2018-12-17
期刊: Microbiome
影响因子: 15.5
作者:
Davis, Nicole M;Proctor, Diana M;Callahan, Benjamin J
通讯作者: Callahan, Benjamin J
DOI: 10.1186/s13059-016-0918-z
发表时间: 2016-03-31
期刊: Genome biology
影响因子: 12.3
作者:
Peltzer A;Jäger G;Herbig A;Seitz A;Kniep C;Krause J;Nieselt K
通讯作者: Nieselt K