Statistical methods for detecting differentially abundant features in clinical metagenomic samples.

Statistical methods for detecting differentially abundant features in clinical metagenomic samples.
复制标题

DOI:
10.1371/journal.pcbi.1000352
复制
发表时间:
2009-04
影响因子:
4.3
通讯作者:
Pop M
Pop M
中科院分区:
生物学2区
文献类型:
--
作者:
White JR;Nagarajan N;Pop M

文献摘要

参考文献

被引文献

相似文献

许多研究目前正在进行中,以表征居住在我们世界的微生物群落。这些研究旨在极大地扩展我们对微生物生物圈的理解,更重要的是,希望揭示我们与肠道细菌菌群之间复杂共生关系的秘密。这些发现的一个重要先决条件是计算工具能够快速准确地比较复杂细菌群落生成的大型数据集,以识别区分它们的特征。我们提出了一种统计方法,用于基于计数数据(例如通过测序获得的)比较来自两个治疗群体的临床宏基因组样品,以检测差异丰富的特征。我们的方法,Metastats,采用错误发现率,以提高在高复杂性环境中的特异性,并单独处理稀疏采样的功能,使用Fisher的精确检验。根据各种模拟,我们表明,Metastats表现良好,以前使用的方法相比,显着优于其他方法的稀疏计数的功能。我们证明了我们的方法在几个数据集上的实用性,包括肥胖和瘦人类肠道微生物组的16S rRNA调查,婴儿和成熟肠道微生物组的COG功能概况,以及从85个宏基因组的随机测序推断的细菌和病毒代谢子系统数据。我们的方法应用于肥胖数据集揭示了原始研究中未报告的肥胖和瘦受试者之间的差异。对于COG和子系统数据集,我们首次对这些人群之间的差异进行了严格的统计评估。本文中描述的方法是第一个解决包括来自多个受试者的样品的临床宏基因组数据集的方法。我们的方法在不同复杂度和采样水平的数据集上都是鲁棒的。虽然专为宏基因组应用而设计,但我们的软件也可应用于数字基因表达研究(例如SAGE)。我们的方法的Web服务器实现和免费的源代码可以在http://metastats.cbcb.umd.edu/上找到。新兴的宏基因组学领域旨在仅通过DNA分析来了解微生物群落的结构和功能。目前比较社区的宏基因组学研究类似于来自两个一般人群(例如生病和健康)的多个受试者的大规模临床试验。为了改进对这类实验数据的分析,我们开发了一种统计方法来检测微生物群落之间的差异丰富特征,即在一个种群中富集或耗尽的特征。我们表明,我们的方法适用于各种宏基因组数据,从分类信息到功能注释。我们还评估了瘦和肥胖人群之间肠道微生物群的分类学差异,以及成熟和婴儿肠道微生物群功能能力之间的差异,以及微生物和病毒宏基因组的差异。我们的方法是第一个在多个受试者的比较宏基因组学研究中统计解决差异丰度的方法,我们希望能让研究人员更全面地了解两种环境的差异。
Numerous studies are currently underway to characterize the microbial communities inhabiting our world. These studies aim to dramatically expand our understanding of the microbial biosphere and, more importantly, hope to reveal the secrets of the complex symbiotic relationship between us and our commensal bacterial microflora. An important prerequisite for such discoveries are computational tools that are able to rapidly and accurately compare large datasets generated from complex bacterial communities to identify features that distinguish them. We present a statistical method for comparing clinical metagenomic samples from two treatment populations on the basis of count data (e.g. as obtained through sequencing) to detect differentially abundant features. Our method, Metastats, employs the false discovery rate to improve specificity in high-complexity environments, and separately handles sparsely-sampled features using Fisher's exact test. Under a variety of simulations, we show that Metastats performs well compared to previously used methods, and significantly outperforms other methods for features with sparse counts. We demonstrate the utility of our method on several datasets including a 16S rRNA survey of obese and lean human gut microbiomes, COG functional profiles of infant and mature gut microbiomes, and bacterial and viral metabolic subsystem data inferred from random sequencing of 85 metagenomes. The application of our method to the obesity dataset reveals differences between obese and lean subjects not reported in the original study. For the COG and subsystem datasets, we provide the first statistically rigorous assessment of the differences between these populations. The methods described in this paper are the first to address clinical metagenomic datasets comprising samples from multiple subjects. Our methods are robust across datasets of varied complexity and sampling level. While designed for metagenomic applications, our software can also be applied to digital gene expression studies (e.g. SAGE). A web server implementation of our methods and freely available source code can be found at http://metastats.cbcb.umd.edu/. The emerging field of metagenomics aims to understand the structure and function of microbial communities solely through DNA analysis. Current metagenomics studies comparing communities resemble large-scale clinical trials with multiple subjects from two general populations (e.g. sick and healthy). To improve analyses of this type of experimental data, we developed a statistical methodology for detecting differentially abundant features between microbial communities, that is, features that are enriched or depleted in one population versus another. We show our methods are applicable to various metagenomic data ranging from taxonomic information to functional annotations. We also provide an assessment of taxonomic differences in gut microbiota between lean and obese humans, as well as differences between the functional capacities of mature and infant gut microbiomes, and those of microbial and viral metagenomes. Our methods are the first to statistically address differential abundance in comparative metagenomics studies with multiple subjects, and we hope will give researchers a more complete picture of how exactly two environments differ.
DOI: 10.1186/1471-2105-7-162
发表时间: 2006-03-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Rodriguez-Brito B;Rohwer F;Edwards RA
通讯作者: Edwards RA
DOI: 10.1126/science.1124234
发表时间: 2006-06-02
期刊: SCIENCE
影响因子: 56.9
作者:
Gill, Steven R.;Pop, Mihai;Nelson, Karen E.
通讯作者: Nelson, Karen E.
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1007/bf00162526
发表时间: 1996-06-01
影响因子: 2.2
作者:
Hesterberg, T
通讯作者: Hesterberg, T
DOI: 10.2307/2289294
发表时间: 1988-09-01
影响因子: 3.7
作者:
JOHNS, MV
通讯作者: JOHNS, MV