pime: A package for discovery of novel differences among microbial communities

pime: A package for discovery of novel differences among microbial communities
复制标题

DOI:
10.1111/1755-0998.13116
复制
发表时间:
2019-12-02
影响因子:
7.7
通讯作者:
Triplett, Eric W.
Triplett, Eric W.
中科院分区:
生物学1区
文献类型:
--
作者:
Roesch, Luiz Fernando W.;Dobbler, Priscila T.;Triplett, Eric W.

文献摘要

被引文献

相似文献

用于分析微生物群落的数据通常是稀疏的,一些微生物在少数样品中具有高丰度,而在其他样品中几乎不存在。然而,目前缺乏能够处理这种稀疏性的生物信息学工具。pime(微生物组评估流行区间)的设计是为了去除那些可能在少数样品中相对丰度高但总体流行率低的分类群。将pime与现有方法进行了可靠性和鲁棒性比较,并使用16S rRNA独立数据集进行了测试。在每个处理流行率区间内未共享的Pime过滤器微生物分类群从5%的流行率开始,每个过滤步骤增加5%。对于每个流行区间,计算了数百个决策树来预测检测治疗差异的可能性。用户通过选择在数据集中保留大部分16S rRNA序列同时显示最低错误率的流行区间来选择最佳流行过滤数据集。为了获得在构建流行过滤数据集时引入类型I错误的可能性,还包括基于错误检测步骤的方法。对已发表数据集的二次分析发现了比以前报道的其他预期的微生物关联,当只考虑相对丰度时,这些关联可能被掩盖。
The data used for profiling microbial communities is usually sparse with some microbes having high abundance in a few samples and being nearly absent in others. However, current bioinformatics tools able to deal with this sparsity are lacking. pime (Prevalence Interval for Microbiome Evaluation) was designed to remove those taxa that may be high in relative abundance in just a few samples but have a low prevalence overall. The reliability and robustness of pime were compared against existing methods and tested using 16S rRNA independent data sets. pime filters microbial taxa not shared in a per treatment prevalence interval started at 5% prevalence with increasing increments of 5% at each filtering step. For each prevalence interval, hundreds of decision trees were calculated to predict the likelihood of detecting differences in treatments. The best prevalence-filtered data set was user-selected by choosing the prevalence interval that kept a large portion of the 16S rRNA sequences in the data set while also showing the lowest error rate. To obtain the likelihood of introducing type I error while building prevalence-filtered data sets, an error detection step based was also included. A pime reanalysis of published data sets uncovered other expected microbial associations than previously reported, which may be masked when only relative abundance was considered.