Machine Learning Meta-analysis of Large Metagenomic Datasets: Tools and Biological Insights.

Machine Learning Meta-analysis of Large Metagenomic Datasets: Tools and Biological Insights.
复制标题

DOI:
10.1371/journal.pcbi.1004977
复制
发表时间:
2016-07
影响因子:
4.3
通讯作者:
Segata N
Segata N
中科院分区:
生物学2区
文献类型:
--
作者:
Pasolli E;Truong DT;Malik F;Waldron L;Segata N

文献摘要

被引文献

相似文献

人类相关微生物组的散弹枪宏基因组分析为人类疾病和健康状况的预测和生物标志物发现提供了丰富的微生物特征。然而,这种高分辨率微生物特征的使用带来了新的挑战,并且缺乏用于学习任务的有效计算工具。此外,分类规则几乎没有在独立研究中得到验证,这对疾病预测模型的普遍性和泛化提出了质疑。在本文中,我们全面评估了基于宏基因组学的预测任务和潜在微生物组表型关联强度的定量评估方法。我们开发了一个计算框架,用于使用定量微生物组谱预测任务,包括物种水平的相对丰度和菌株特异性标记的存在。对来自8项大规模研究的2424个公开可获得的宏基因组样本进行了全面的荟萃分析,特别强调了跨队列的泛化。交叉验证显示了良好的疾病预测能力,通常通过特征选择和使用菌株特异性标记而不是物种水平的分类丰度来提高疾病预测能力。在交叉研究分析中,在研究之间转移的模型在某些情况下比研究内交叉验证检验的模型更不准确。有趣的是,将来自其他研究的健康(对照)样本添加到训练集提高了疾病预测能力。一些微生物物种(最明显的是血管链球菌)似乎表征了微生物群的一般生态失调状态,而不是与特定疾病的联系。我们在模拟“健康”微生物组特征方面的结果可以被认为是定义一般微生物生态失调的第一步。软件框架、微生物组配置文件和数千个样本的元数据可在http://segatalab.cibio.unitn.it/tools/metaml上公开获取。人类微生物群——与人类宿主相关的全部微生物——与宿主的免疫和代谢功能密切相互作用,对人类健康至关重要。通过下一代DNA测序技术,在与健康和患病个体相关的微生物组表征方面取得了重大进展,该技术可以直接从未培养的人类相关样品(例如粪便)中准确估计微生物群落。特别是,霰弹枪宏基因组学提供了前所未有的物种和品系分辨率水平的数据。一些大规模的宏基因组疾病相关数据集也正在变得可用,并且已经提出了基于宏基因组特征的疾病预测模型。然而,对不同人群和疾病的预测模型的泛化尚未得到验证。在本文中,我们全面评估了基于宏基因组学的预测任务和微生物组表型关联的定量评估方法。我们考虑了来自8项研究和6种不同疾病的2424个样本,以评估基于散弹枪宏基因组数据建立的模型的独立预测准确性,并比较微生物组作为预测工具的实际使用策略。
Shotgun metagenomic analysis of the human associated microbiome provides a rich set of microbial features for prediction and biomarker discovery in the context of human diseases and health conditions. However, the use of such high-resolution microbial features presents new challenges, and validated computational tools for learning tasks are lacking. Moreover, classification rules have scarcely been validated in independent studies, posing questions about the generality and generalization of disease-predictive models across cohorts. In this paper, we comprehensively assess approaches to metagenomics-based prediction tasks and for quantitative assessment of the strength of potential microbiome-phenotype associations. We develop a computational framework for prediction tasks using quantitative microbiome profiles, including species-level relative abundances and presence of strain-specific markers. A comprehensive meta-analysis, with particular emphasis on generalization across cohorts, was performed in a collection of 2424 publicly available metagenomic samples from eight large-scale studies. Cross-validation revealed good disease-prediction capabilities, which were in general improved by feature selection and use of strain-specific markers instead of species-level taxonomic abundance. In cross-study analysis, models transferred between studies were in some cases less accurate than models tested by within-study cross-validation. Interestingly, the addition of healthy (control) samples from other studies to training sets improved disease prediction capabilities. Some microbial species (most notably Streptococcus anginosus) seem to characterize general dysbiotic states of the microbiome rather than connections with a specific disease. Our results in modelling features of the “healthy” microbiome can be considered a first step toward defining general microbial dysbiosis. The software framework, microbiome profiles, and metadata for thousands of samples are publicly available at http://segatalab.cibio.unitn.it/tools/metaml. The human microbiome–the entire set of microbial organisms associated with the human host–interacts closely with host immune and metabolic functions and is crucial for human health. Significant advances in the characterization of the microbiome associated with healthy and diseased individuals have been obtained through next-generation DNA sequencing technologies, which permit accurate estimation of microbial communities directly from uncultured human-associated samples (e.g., stool). In particular, shotgun metagenomics provide data at unprecedented species- and strain- levels of resolution. Several large-scale metagenomic disease-associated datasets are also becoming available, and disease-predictive models built on metagenomic signatures have been proposed. However, the generalization of resulting prediction models on different cohorts and diseases has not been validated. In this paper, we comprehensively assess approaches to metagenomics-based prediction tasks and for quantitative assessment of microbiome-phenotype associations. We consider 2424 samples from eight studies and six different diseases to assess the independent prediction accuracy of models built on shotgun metagenomic data and to compare strategies for practical use of the microbiome as a prediction tool.