Normalization and microbial differential abundance strategies depend upon data characteristics.

Normalization and microbial differential abundance strategies depend upon data characteristics.
复制标题

DOI:
10.1186/s40168-017-0237-y
复制
发表时间:
2017-03-03
期刊:
影响因子:
15.5
通讯作者:
Knight R
Knight R
中科院分区:
生物学1区
文献类型:
--
作者:
Weiss S;Xu ZZ;Peddada S;Amir A;Bittinger K;Gonzalez A;Lozupone C;Zaneveld JR;Vázquez-Baeza Y;Birmingham A;Hyde ER;Knight R

文献摘要

被引文献

相似文献

来自16 S核糖体RNA(rRNA)扩增子测序的数据对生态学和统计学解释提出了挑战。特别是,库大小通常在几个量级范围内变化,并且数据包含许多零。虽然我们通常对比较两个或多个类群的生态系统中分类群的相对丰度感兴趣,但我们只能测量从生态系统中获得的标本中的分类群相对丰度。由于标本中分类单元相对多度的比较并不等同于生态系统中分类单元相对多度的比较,这就提出了一个特殊的挑战。第二,因为标本中(以及生态系统中)类群的相对丰度总和为1,所以这些是组成数据。由于成分数据受单纯形(和为1)的约束,并且在欧几里得空间中不是无约束的,因此许多标准的分析方法都不适用。在这里,我们评估这些挑战如何影响现有的归一化方法和差异丰度分析的性能。对正常化的影响:大多数标准化方法能够根据生物来源成功地聚类样品时,组在其总体微生物组成有很大差异。稀疏更清楚地聚类样品根据生物起源比其他标准化技术做排序指标的基础上存在或不存在。由于库的大小,替代的标准化措施可能容易受到伪影的影响。对差异丰度测试的影响:我们建立在以前的工作,以评估七个建议的统计方法,使用稀薄以及原始数据。我们的模拟研究表明,许多差异丰度测试方法的错误发现率并没有增加稀疏本身,虽然当然稀疏的结果在灵敏度的损失,由于消除了一部分可用的数据。对于平均文库大小差异较大(约10倍)的组,稀疏化降低了错误发现率。DESeq 2在没有添加常数的情况下,在较小的数据集(每组<20个样品)上增加了灵敏度,但在更多的样品、非常不均匀(~10×)的文库大小和/或组成效应的情况下倾向于更高的错误发现率。为了推断生态系统中的类群丰度,微生物组组成分析(ANCOM)不仅非常敏感(每组>20个样本),而且是唯一一种可以很好地控制错误发现率的方法。这些发现指导了根据给定研究的数据特征使用哪种标准化和差异丰度技术。本文的在线版本(doi:10.1186/s40168-017-0237-y)包含补充材料,可供授权用户使用。
Data from 16S ribosomal RNA (rRNA) amplicon sequencing present challenges to ecological and statistical interpretation. In particular, library sizes often vary over several ranges of magnitude, and the data contains many zeros. Although we are typically interested in comparing relative abundance of taxa in the ecosystem of two or more groups, we can only measure the taxon relative abundance in specimens obtained from the ecosystems. Because the comparison of taxon relative abundance in the specimen is not equivalent to the comparison of taxon relative abundance in the ecosystems, this presents a special challenge. Second, because the relative abundance of taxa in the specimen (as well as in the ecosystem) sum to 1, these are compositional data. Because the compositional data are constrained by the simplex (sum to 1) and are not unconstrained in the Euclidean space, many standard methods of analysis are not applicable. Here, we evaluate how these challenges impact the performance of existing normalization methods and differential abundance analyses. Effects on normalization: Most normalization methods enable successful clustering of samples according to biological origin when the groups differ substantially in their overall microbial composition. Rarefying more clearly clusters samples according to biological origin than other normalization techniques do for ordination metrics based on presence or absence. Alternate normalization measures are potentially vulnerable to artifacts due to library size. Effects on differential abundance testing: We build on a previous work to evaluate seven proposed statistical methods using rarefied as well as raw data. Our simulation studies suggest that the false discovery rates of many differential abundance-testing methods are not increased by rarefying itself, although of course rarefying results in a loss of sensitivity due to elimination of a portion of available data. For groups with large (~10×) differences in the average library size, rarefying lowers the false discovery rate. DESeq2, without addition of a constant, increased sensitivity on smaller datasets (<20 samples per group) but tends towards a higher false discovery rate with more samples, very uneven (~10×) library sizes, and/or compositional effects. For drawing inferences regarding taxon abundance in the ecosystem, analysis of composition of microbiomes (ANCOM) is not only very sensitive (for >20 samples per group) but also critically the only method tested that has a good control of false discovery rate. These findings guide which normalization and differential abundance techniques to use based on the data characteristics of a given study. The online version of this article (doi:10.1186/s40168-017-0237-y) contains supplementary material, which is available to authorized users.