Variational Approximation-Based Model Selection for Microbial Network Inference

Variational Approximation-Based Model Selection for Microbial Network Inference
复制标题

DOI:
10.1089/cmb.2021.0595
复制
发表时间:
2022-05
期刊:
Journal of computational biology : a journal of computational molecular cell biology
影响因子:
--
通讯作者:
Shibu Yooseph;Sahar Tavakoli
Shibu Yooseph;Sahar Tavakoli
中科院分区:
其他
文献类型:
--
作者:
Shibu Yooseph;Sahar Tavakoli

文献摘要

相似文献

微生物关联是指微生物群落中各组成类群之间直接或间接的相互作用,对群落的结构、组织和功能起着重要的决定作用。微生物关联可以用加权图(微生物网络)表示,其节点表示分类群,边表示成对关联。微生物网络通常是从样品-分类群矩阵中推断出来的,该矩阵是通过对多个生物样品进行测序并确定每个样品中的分类群计数而获得的。然而,众所周知,微生物关联受到环境和/或宿主因素的影响。因此,微生物组研究中产生的样本-分类群矩阵涉及环境和/或临床元数据变量的广泛值,实际上可能与多个微生物网络相关联。在这项研究中,我们考虑了从给定的样本-分类群计数矩阵推断多个微生物网络的问题。每个样本是一个计数向量,假设是由多元泊松对数正态分布组成的混合模型生成的。针对模型选择问题,提出了一种变分期望最大化算法来推断该混合模型的正确分量数。我们的方法包括将混合模型重新构建为潜在变量模型,仅将混合系数作为参数,然后使用证据下界框架近似边际似然。我们的算法在使用不同图结构(带、集线器、聚类、随机和无标度)的集合生成的大型模拟数据集上进行评估。
Microbial associations are characterized by both direct and indirect interactions between the constituent taxa in a microbial community, and play an important role in determining the structure, organization, and function of the community. Microbial associations can be represented using a weighted graph (microbial network), whose nodes represent taxa and edges represent pairwise associations. A microbial network is typically inferred from a sample-taxa matrix that is obtained by sequencing multiple biological samples and identifying the taxa counts in each sample. However, it is known that microbial associations are impacted by environmental and/or host factors. Thus, a sample-taxa matrix generated in a microbiome study involving a wide range of values for the environmental and/or clinical metadata variables may in fact be associated with more than one microbial network. In this study, we consider the problem of inferring multiple microbial networks from a given sample-taxa count matrix. Each sample is a count vector assumed to be generated by a mixture model consisting of component distributions that are multivariate Poisson log-normal. We present a variational expectation maximization algorithm for the model selection problem to infer the correct number of components of this mixture model. Our approach involves reframing the mixture model as a latent variable model, treating only the mixing coefficients as parameters, and subsequently approximating the marginal likelihood using an evidence lower bound framework. Our algorithm is evaluated on a large simulated dataset generated using a collection of different graph structures (band, hub, cluster, random, and scale-free).