Analysis of composition of microbiomes: a novel method for studying microbial composition.

Analysis of composition of microbiomes: a novel method for studying microbial composition.
复制标题

DOI:
10.3402/mehd.v26.27663
复制
发表时间:
2015
影响因子:
--
通讯作者:
Peddada SD
Peddada SD
中科院分区:
其他
文献类型:
--
作者:
Mandal S;Van Treuren W;White RA;Eggesbø M;Knight R;Peddada SD

文献摘要

被引文献

相似文献

理解调节我们微生物群落的因素很重要,但需要合适的统计学方法。在比较两个或多个群体时,大多数现有方法要么忽视微生物组数据中潜在的组成结构,要么使用诸如多项分布和狄利克雷 - 多项分布等概率模型,这些模型可能强加一种不适合微生物组数据的相关结构。 为了开发一种考虑组成限制的方法,以减少在生态系统层面检测差异丰富的分类群时的错误发现,同时保持较高的统计功效。 我们引入了一种新的统计框架,称为微生物群落组成分析(ANCOM)。ANCOM考虑了数据中的潜在结构,可用于比较两个或多个群体中微生物群落的组成。ANCOM不做分布假设,并且可以在线性模型框架中实施,以调整协变量以及对纵向数据进行建模。ANCOM在比较涉及数千个分类群的样本时也具有良好的扩展性。 我们将ANCOM的性能与标准t检验以及一种最近发表的称为零膨胀高斯(ZIG)方法进行了比较,该方法用于对两个或多个群体中的平均分类群丰度进行推断。ANCOM将错误发现率(FDR)控制在期望的名义水平,同时还提高了功效,而t检验和ZIG的FDR则过高,在某些情况下,t检验高达68%,ZIG高达60%。我们使用两个人类肠道中公开可用的微生物数据集说明了ANCOM的性能,证明了它在检验微生物群落组成差异假设方面的普遍适用性。 使用对数比率分析考虑组成性可显著改善微生物群落调查数据中的推断。
Understanding the factors regulating our microbiota is important but requires appropriate statistical methodology. When comparing two or more populations most existing approaches either discount the underlying compositional structure in the microbiome data or use probability models such as the multinomial and Dirichlet-multinomial distributions, which may impose a correlation structure not suitable for microbiome data. To develop a methodology that accounts for compositional constraints to reduce false discoveries in detecting differentially abundant taxa at an ecosystem level, while maintaining high statistical power. We introduced a novel statistical framework called analysis of composition of microbiomes (ANCOM). ANCOM accounts for the underlying structure in the data and can be used for comparing the composition of microbiomes in two or more populations. ANCOM makes no distributional assumptions and can be implemented in a linear model framework to adjust for covariates as well as model longitudinal data. ANCOM also scales well to compare samples involving thousands of taxa. We compared the performance of ANCOM to the standard t-test and a recently published methodology called Zero Inflated Gaussian (ZIG) methodology for drawing inferences on the mean taxa abundance in two or more populations. ANCOM controlled the false discovery rate (FDR) at the desired nominal level while also improving power, whereas the t-test and ZIG had inflated FDRs, in some instances as high as 68% for the t-test and 60% for ZIG. We illustrate the performance of ANCOM using two publicly available microbial datasets in the human gut, demonstrating its general applicability to testing hypotheses about compositional differences in microbial communities. Accounting for compositionality using log-ratio analysis results in significantly improved inference in microbiota survey data.