A zero inflated log-normal model for inference of sparse microbial association networks.

A zero inflated log-normal model for inference of sparse microbial association networks.
复制标题

零膨胀对数正态模型用于稀疏微生物关联网络的推断。

DOI:
10.1371/journal.pcbi.1009089
复制
发表时间:
2021-06
影响因子:
4.3
通讯作者:
Brüls T
Brüls T
中科院分区:
生物学2区
文献类型:
--
作者:
Prost V;Gazut S;Brüls T

文献摘要

参考文献

被引文献

相似文献

高通量宏基因组测序的出现促进了有效的分类学分析方法的发展,可以测量各种环境样本中生物体的存在、丰度和系统发育。多元序列衍生的丰度数据进一步有可能推断微生物种群之间的生态关联,但需要考虑几个技术问题,例如数据的组成性质、其极度稀疏和过度分散,以及经常需要在不确定的情况下操作。生态网络重建问题经常被纳入高斯图形模型(GGM)的范式中,其中可以使用有效的结构推理算法,例如图形套索和邻域选择。不幸的是,GGM 或其变体无法正确解释现实世界宏基因组分类谱中出现的极其稀疏的模式。特别是,大多数统计方法无法正确处理与生物信号的真实缺失相对应的结构零点(与采样零点相对)。我们在这里提出了一个零膨胀对数正态图形模型(可在 https://github.com/vincentprost/Zi-LN 获得),专门用于处理此类“生物”零,并展示了与用于推断微生物关联网络的最先进统计方法相比的显着性能增益,其中在分析显示与真实世界宏基因组数据集相同的稀疏水平的分类资料时获得了最显着的增益。协会在社区成员的结构和动态中的重要性得到了广泛认可,但我们目前无法对从环境中采样的大多数微生物进行共培养。因此,预测微生物关联的计算方法具有实际意义,特别是考虑到宏基因组学生成的大量多元微生物丰度数据。理论上,这些数据可以用来推断关联网络,但迄今为止取得的成功有限,因为它的一些属性导致了技术困难,包括其极端稀疏性、组合性和过度分散等。特别是,与生物信号的真实缺失相对应的结构零(与采样和技术零相对)经常无法正确处理,并且这种非随机缺失可能导致高水平的误报。鉴于零值的普遍性,建模过程应该首先考虑零生成过程来正确处理零值。我们在这里描述了一种截断的对数正态图形模型,该模型专门解决了源自生物缺失的零,并讨论了估计稀疏和高维关联网络的一致方法。我们还表明,该模型生成的稀疏多元计数更接近来自现实世界微生物组的计数。
The advent of high-throughput metagenomic sequencing has prompted the development of efficient taxonomic profiling methods allowing to measure the presence, abundance and phylogeny of organisms in a wide range of environmental samples. Multivariate sequence-derived abundance data further has the potential to enable inference of ecological associations between microbial populations, but several technical issues need to be accounted for, like the compositional nature of the data, its extreme sparsity and overdispersion, as well as the frequent need to operate in under-determined regimes. The ecological network reconstruction problem is frequently cast into the paradigm of Gaussian Graphical Models (GGMs) for which efficient structure inference algorithms are available, like the graphical lasso and neighborhood selection. Unfortunately, GGMs or variants thereof can not properly account for the extremely sparse patterns occurring in real-world metagenomic taxonomic profiles. In particular, structural zeros (as opposed to sampling zeros) corresponding to true absences of biological signals fail to be properly handled by most statistical methods. We present here a zero-inflated log-normal graphical model (available at https://github.com/vincentprost/Zi-LN) specifically aimed at handling such “biological” zeros, and demonstrate significant performance gains over state-of-the-art statistical methods for the inference of microbial association networks, with most notable gains obtained when analyzing taxonomic profiles displaying sparsity levels on par with real-world metagenomic datasets. The importance of associations in the structuring and dynamics of community members is widely acknowledged, but we are currently unable to co-culture most of the micro-organims sampled from the environment. Computational methods to predict microbial associations can therefore be of practical interest, in particular given the large amounts of multivariate microbial abundance data generated by metagenomics. This data can in theory be leveraged to infer association networks, but with limited success so far, as several of its attributes lead to technical difficulties, including its extreme sparsity, compositionality and overdispersion among others. In particular, structural zeros (as opposed to sampling and technical zeros) corresponding to true absences of biological signals frequently fail to be properly handled, and such non-random absences can lead to high levels of false positives. Given their prevalence, zero values should be properly handled by the modeling process by accounting for the zero generating process in the first place. We describe here a truncated log-normal graphical model that specifically addresses zeros originating from biological absences, and discuss consistent methods for estimating sparse and high-dimensional association networks. We also show that this model generates sparse multivariate counts more close to those derived from real-world microbiomes.
DOI: 10.1371/journal.pcbi.1002606
发表时间: 2012
影响因子: 4.3
作者:
Faust K;Sathirapongsasuti JF;Izard J;Segata N;Gevers D;Raes J;Huttenhower C
通讯作者: Huttenhower C
DOI: 10.1038/s41467-020-18127-y
发表时间: 2020-08-28
影响因子: 16.6
作者:
Ianiro, Gianluca;Rossi, Ernesto;Cammarota, Giovanni
通讯作者: Cammarota, Giovanni
DOI: 10.1371/journal.pcbi.1005852
发表时间: 2017-11
影响因子: 4.3
作者:
Schwager E;Mallick H;Ventz S;Huttenhower C
通讯作者: Huttenhower C
DOI: 10.1089/cmb.2016.0061
发表时间: 2016-06-01
影响因子: 1.7
作者:
Biswas, Surojit;Mcdonald, Meredith;Jojic, Vladimir
通讯作者: Jojic, Vladimir
DOI: 10.1186/s12859-019-2882-6
发表时间: 2019-11-22
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Patuzzi, Ilaria;Baruzzo, Giacomo;Di Camillo, Barbara
通讯作者: Di Camillo, Barbara