Improved normalization of species count data in ecology by scaling with ranked subsampling (SRS): application to microbial communities

Improved normalization of species count data in ecology by scaling with ranked subsampling (SRS): application to microbial communities
复制标题

DOI:
10.7717/peerj.9593
复制
发表时间:
2020-08-03
期刊:
影响因子:
2.7
通讯作者:
Karlovsky, Petr
Karlovsky, Petr
中科院分区:
生物学3区
文献类型:
--
作者:
Beule, Lukas;Karlovsky, Petr

文献摘要

被引文献

相似文献

背景生态学中物种计数数据的分析通常需要归一化到相同的样本大小。稀疏化(没有替换的随机子采样)是目前标准的归一化方法,由于其重现性差和潜在的群落结构扭曲而受到广泛批评。在微生物组计数数据的背景下,研究人员明确建议不要使用稀疏。在这里,我们介绍了一种标准化方法的物种计数数据称为缩放与排序子采样(SRS),并证明其适用于微生物群落的分析。SRS包括两个步骤。在缩放步骤中,将所有物种或操作分类单元(OTU)的计数除以缩放因子,所述缩放因子以使得缩放计数的总和等于所选择的计数总数C-min的方式选择。所有OTU的相对频率保持不变。在随后的分级二次采样步骤中,通过算法将非整数计数值转换为整数,该算法最小化关于种群结构(物种或OTU的相对频率)的二次采样误差,同时保持计数总数等于C-minSRS和稀疏化通过归一化代表土壤细菌群落的测试文库进行比较。对于通过稀疏化归一化为不同大小的文库以及每个具有10,000个重复的SRS,确定生物多样性和种群结构的共同参数(香农指数H ',物种丰富度,物种组成和OTU的相对丰度)。SRS在R中的实现可供下载(https://doi.org/10.20387/BONARES-2657-1NP 3)。SRS表现出更大的可重复性和保存OTU频率和α多样性比稀疏。Shannon多样性的方差随着稀疏化后文库大小的减小而增加,但SRS保持为零。OTU的相对丰度在通过稀疏化产生的文库之间变化很大,而通过SRS标准化的文库仅显示可忽略的变化。通过稀疏化标准化的相同文库的重复之间的Bray-Curtis相异性指数揭示了物种组成的大的变化,其在一些稀疏到小尺寸的文库中达到完全相异性(没有一个OTU共享)。通过SRS归一化的复制文库之间的差异在每个文库大小下保持可忽略不计的低。在稀疏化后,差异性的方差随着文库大小的减小而增加,而在SRS后,差异性的方差保持为零或可忽略不计。通过分级二次采样缩放的OTU或物种计数的归一化通过最小化二次采样误差来保留原始群落结构。因此,我们提出SRS的生物计数数据的标准化。
Background. Analysis of species count data in ecology often requires normalization to an identical sample size. Rarefying (random subsampling without replacement), which is the current standard method for normalization, has been widely criticized for its poor reproducibility and potential distortion of the community structure. In the context of microbiome count data, researchers explicitly advised against the use of rarefying. Here we introduce a normalization method for species count data called scaling with ranked subsampling (SRS) and demonstrate its suitability for the analysis of microbial communities.Methods. SRS consists of two steps. In the scaling step, the counts for all species or operational taxonomic units (OTUs) are divided by a scaling factor chosen in such a way that the sum of scaled counts equals the selected total number of counts C-min. The relative frequencies of all OTUs remain unchanged. In the subsequent ranked subsampling step, non-integer count values are converted into integers by an algorithm that minimizes subsampling error with regard to the population structure (relative frequencies of species or OTUs) while keeping the total number of counts equal C-min. SRS and rarefying were compared by normalizing a test library representing a soil bacterial community. Common parameters of biodiversity and population structure (Shannon index H', species richness, species composition, and relative abundances of OTUs) were determined for libraries normalized to different size by rarefying as well as SRS with 10,000 replications each. An implementation of SRS in R is available for download (https://doi.org/10.20387/BONARES-2657-1NP3).Results. SRS showed greater reproducibility and preserved OTU frequencies and alpha diversity better than rarefying. The variance in Shannon diversity increased with the reduction of the library size after rarefying but remained zero for SRS. Relative abundances of OTUs strongly varied among libraries generated by rarefying, whereas libraries normalized by SRS showed only negligible variation. Bray-Curtis index of dissimilarity among replicates of the same library normalized by rarefying revealed a large variation in species composition, which reached complete dissimilarity (not a single OTU shared) among some libraries rarefied to a small size. The dissimilarity among replicated libraries normalized by SRS remained negligibly low at each library size. The variance in dissimilarity increased with the decreasing library size after rarefying, whereas it remained either zero or negligibly low after SRS.Conclusions. Normalization of OTU or species counts by scaling with ranked subsampling preserves the original community structure by minimizing subsampling errors. We therefore propose SRS for the normalization of biological count data.