Microbiome, Metagenomics, and High-Dimensional Compositional Data Analysis

Microbiome, Metagenomics, and High-Dimensional Compositional Data Analysis
复制标题

DOI:
10.1146/annurev-statistics-010814-020351
复制
发表时间:
2015-01-01
期刊:
ANNUAL REVIEW OF STATISTICS AND ITS APPLICATION, VOL 2
影响因子:
--
通讯作者:
Li, Hongzhe
Li, Hongzhe
中科院分区:
其他
文献类型:
--
作者:
Li, Hongzhe

文献摘要

被引文献

相似文献

人体微生物群是人体内和体内所有微生物的总和,其在健康和疾病中的重要性日益被认识到。高通量测序技术最近使科学家能够对构成微生物组的所有微生物进行公正的量化。通常,一个样本可以产生数以亿计的短测序读数。然而,新技术产生的数据的独特特征,以及这些数据的巨大规模,使得从微生物组研究中得出有效的生物学推断变得困难。对这些大数据的分析带来了巨大的统计和计算挑战。重要的问题包括相对分类群、细菌基因和代谢丰度的标准化和量化;将系统发育信息纳入元基因组数据分析;以及高维成分数据的多变量分析。我们回顾了现有的方法,指出了它们的局限性,并概述了未来的研究方向。
The human microbiome is the totality of all microbes in and on the human body, and its importance in health and disease has been increasingly recognized. High-throughput sequencing technologies have recently enabled scientists to obtain an unbiased quantification of all microbes constituting the microbiome. Often, a single sample can produce hundreds of millions of short sequencing reads. However, unique characteristics of the data produced by the new technologies, as well as the sheer magnitude of these data, make drawing valid biological inferences from microbiome studies difficult. Analysis of these big data poses great statistical and computational challenges. Important issues include normalization and quantification of relative taxa, bacterial genes, and metabolic abundances; incorporation of phylogenetic information into analysis of metagenomics data; and multivariate analysis of high-dimensional compositional data. We review existing methods, point out their limitations, and outline future research directions.