Evaluating supervised and unsupervised background noise correction in human gut microbiome data.

Evaluating supervised and unsupervised background noise correction in human gut microbiome data.
复制标题

DOI:
10.1371/journal.pcbi.1009838
复制
发表时间:
2022-03
影响因子:
4.3
通讯作者:
Garud NR
Garud NR
中科院分区:
生物学2区
文献类型:
--
作者:
Briscoe L;Balliu B;Sankararaman S;Halperin E;Garud NR

文献摘要

参考文献

被引文献

相似文献

从宏基因组数据中预测人类表型和识别疾病生物标志物的能力对于开发微生物组相关疾病的治疗方法至关重要。然而,宏基因组数据通常受到与感兴趣的表型无关的技术变量的影响,例如测序方案,这可能使预测表型和发现疾病的生物标志物变得困难。纠正背景噪声的监督方法最初是为基因表达和RNA-seq数据设计的,通常应用于微生物组数据,但可能受到限制,因为它们不能解释未测量的变异来源。无监督方法解决了这一问题,但目前的方法是有限的,因为它们不具备处理微生物组数据的独特方面的能力,这些数据是组成的,高度倾斜的,稀疏的。我们对不同去噪变换与监督校正方法以及目前在其他领域使用但尚未应用于微生物组数据的无监督主成分校正方法相结合的能力进行了比较分析。我们发现,无监督主成分校正方法在减少生物标志物的错误发现方面具有与监督方法相当的能力,并且不需要知道先验变异的来源。然而,在预测任务中,只有当技术变量对数据中的大部分方差做出贡献时,它似乎才能改善预测。随着新的和更大的宏基因组数据集越来越多,背景噪声校正将成为生成可重复微生物组分析的必要条件。众所周知,人类肠道微生物群在健康中发挥着重要作用,并与许多疾病有关,包括结肠直肠癌、肥胖和糖尿病。宿主表型的预测和疾病生物标志物的鉴定对于利用微生物组的治疗潜力至关重要。然而,许多宏基因组数据集受到技术变量的影响,这些技术变量引入了不必要的变异,从而混淆了预测表型和识别生物标志物的能力。目前,最初为基因表达和RNA-seq数据设计的监督方法通常应用于微生物组数据以校正背景噪声,但它们的局限性在于它们不能校正未测量的变异源。无监督方法解决了这一问题,但目前的方法是有限的,因为它们不具备处理微生物组数据的独特方面的能力,这些数据是组成的,高度倾斜的,稀疏的。我们对不同去噪变换与监督校正方法以及无监督主成分校正方法相结合的能力进行了比较分析,发现所有校正方法都减少了生物标志物发现的假阳性。在预测表型的任务中,不同的方法有不同的成功,当技术变量导致数据中的大部分方差时,无监督校正可以改善预测。随着新的和更大的宏基因组数据集越来越多,背景噪声校正将成为生成可重复微生物组分析的必要条件。
The ability to predict human phenotypes and identify biomarkers of disease from metagenomic data is crucial for the development of therapeutics for microbiome-associated diseases. However, metagenomic data is commonly affected by technical variables unrelated to the phenotype of interest, such as sequencing protocol, which can make it difficult to predict phenotype and find biomarkers of disease. Supervised methods to correct for background noise, originally designed for gene expression and RNA-seq data, are commonly applied to microbiome data but may be limited because they cannot account for unmeasured sources of variation. Unsupervised approaches address this issue, but current methods are limited because they are ill-equipped to deal with the unique aspects of microbiome data, which is compositional, highly skewed, and sparse. We perform a comparative analysis of the ability of different denoising transformations in combination with supervised correction methods as well as an unsupervised principal component correction approach that is presently used in other domains but has not been applied to microbiome data to date. We find that the unsupervised principal component correction approach has comparable ability in reducing false discovery of biomarkers as the supervised approaches, with the added benefit of not needing to know the sources of variation apriori. However, in prediction tasks, it appears to only improve prediction when technical variables contribute to the majority of variance in the data. As new and larger metagenomic datasets become increasingly available, background noise correction will become essential for generating reproducible microbiome analyses. The human gut microbiome is known to play a major role in health and is associated with many diseases including colorectal cancer, obesity, and diabetes. The prediction of host phenotypes and identification of biomarkers of disease is essential for harnessing the therapeutic potential of the microbiome. However, many metagenomic datasets are affected by technical variables that introduce unwanted variation that can confound the ability to predict phenotypes and identify biomarkers. Currently, supervised methods originally designed for gene expression and RNA-seq data are commonly applied to microbiome data for correction of background noise, but they are limited in that they cannot correct for unmeasured sources of variation. Unsupervised approaches address this issue, but current methods are limited because they are ill-equipped to deal with the unique aspects of microbiome data, which is compositional, highly skewed, and sparse. We perform a comparative analysis of the ability of different denoising transformations in combination with supervised correction methods as well as an unsupervised principal component correction approach and find that all correction approaches reduce false positives for biomarker discovery. In the task of predicting phenotypes, different approaches have varying success where the unsupervised correction can improve prediction when technical variables contribute to the majority of variance in the data. As new and larger metagenomic datasets become increasingly available, background noise correction will become essential for generating reproducible microbiome analyses.
DOI: 10.1038/s41467-017-01973-8
发表时间: 2017-12-05
影响因子: 16.6
作者:
Duvallet C;Gibbons SM;Gurry T;Irizarry RA;Alm EJ
通讯作者: Alm EJ
DOI: 10.1186/s12866-015-0351-6
发表时间: 2015-03-21
期刊: BMC microbiology
影响因子: 4.2
作者:
Brooks JP;Edwards DJ;Harwich MD Jr;Rivera MC;Fettweis JM;Serrano MG;Reris RA;Sheth NU;Huang B;Girerd P;Vaginal Microbiome Consortium;Strauss JF 3rd;Jefferson KK;Buck GA
通讯作者: Buck GA
DOI: 10.1371/journal.pone.0160169
发表时间: 2016-08-11
期刊: PLOS ONE
影响因子: 3.7
作者:
Cao, Kim-Anh Le;Costello, Mary-Ellen;Rondeau, Pascale
通讯作者: Rondeau, Pascale
DOI: 10.1038/nbt.3960
发表时间: 2017-11-01
影响因子: 46.9
作者:
Costea, Paul I.;Zeller, Georg;Bork, Peer
通讯作者: Bork, Peer
DOI: 10.1186/2049-2618-2-15
发表时间: 2014
期刊: Microbiome
影响因子: 15.5
作者:
Fernandes AD;Reid JN;Macklaim JM;McMurrough TA;Edgell DR;Gloor GB
通讯作者: Gloor GB