Monte Carlo samplers for efficient network inference.

Monte Carlo samplers for efficient network inference.
复制标题

DOI:
10.1371/journal.pcbi.1011256
复制
发表时间:
2023-07
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

在驱动生物过程的底层网络上获取信息通常涉及中断该过程并收集快照数据。当快照数据是随机的时,数据的结构需要概率描述来推断潜在的反应网络。例如,我们可以想象想要从单分子RNA荧光原位杂交(RNA-FISH)中收集的数据类型中学习基因状态网络。在我们考虑的网络中,节点代表网络状态,边代表连接状态的生化反应速率。从快照数据同时估计节点和组成参数的数量仍然是一项具有挑战性的任务,部分原因是数据的不确定性和介导网络的动力学参数之间的时间尺度分离。虽然参数贝叶斯方法在具有严格传播的测量不确定性的网络结构(具有已知节点数量)的情况下学习参数,但学习具有潜在大时间尺度间隔的节点数量和参数仍然是悬而未决的问题。在这里,我们提出了一个贝叶斯非参数框架,并描述了一个混合贝叶斯马尔可夫链蒙特卡罗(MCMC)采样器直接解决这些挑战。特别是,在我们的混合方法中,Hamilton Monte Carlo(HMC)利用局部后验几何推理来探索参数空间;自适应大都会Hastings(AMH)学习合理参数集之间的相关性,以有效地提出可能的模型;并行回火同时考虑多个模型,并具有回火信息内容,以提高采样效率。我们将我们的方法应用于模拟单分子RNA-FISH的合成数据,这是一种探测转录网络的流行快照方法,以说明从这些快照中学习动态模型所固有的挑战以及我们的方法如何解决这些挑战。解码生化反应网络的物种数量,连接它们的反应和反应速率,从快照数据,如单分子荧光原位杂交(smFISH)实验得出的数据,是至关重要的利益。然而,这项任务的挑战目前是解决与网络规范算法,因为:1)网络的大小不能指定独立的速率; 2)速率可能是分开的数量级,生成刚性ODE; 3)不确定性不会传播到网络和参数。我们提出了,最近的计算统计工具的启发,同时推导出反应网络和相关的速率从快照数据,同时传播所有未知的错误的方法。我们通过将网络模型视为随机变量来实现这一点,并且在贝叶斯非参数范式中,开发了模型本身的后验。这种多维后验自然包含多个山丘和山谷,因此我们提出了一种采样器组合,允许对网络及其相关速率进行首次同时和自洽的推断。我们的方法能够处理任意数量的状态,这与目前的技术水平形成了对比,并且可能会修改以前基于网络确定的生物学结论。我们证明了模仿smFISH实验的合成数据的方法,并证明了其对朴素MCMC方案的改进。
Accessing information on an underlying network driving a biological process often involves interrupting the process and collecting snapshot data. When snapshot data are stochastic, the data’s structure necessitates a probabilistic description to infer underlying reaction networks. As an example, we may imagine wanting to learn gene state networks from the type of data collected in single molecule RNA fluorescence in situ hybridization (RNA-FISH). In the networks we consider, nodes represent network states, and edges represent biochemical reaction rates linking states. Simultaneously estimating the number of nodes and constituent parameters from snapshot data remains a challenging task in part on account of data uncertainty and timescale separations between kinetic parameters mediating the network. While parametric Bayesian methods learn parameters given a network structure (with known node numbers) with rigorously propagated measurement uncertainty, learning the number of nodes and parameters with potentially large timescale separations remain open questions. Here, we propose a Bayesian nonparametric framework and describe a hybrid Bayesian Markov Chain Monte Carlo (MCMC) sampler directly addressing these challenges. In particular, in our hybrid method, Hamiltonian Monte Carlo (HMC) leverages local posterior geometries in inference to explore the parameter space; Adaptive Metropolis Hastings (AMH) learns correlations between plausible parameter sets to efficiently propose probable models; and Parallel Tempering takes into account multiple models simultaneously with tempered information content to augment sampling efficiency. We apply our method to synthetic data mimicking single molecule RNA-FISH, a popular snapshot method in probing transcriptional networks to illustrate the identified challenges inherent to learning dynamical models from these snapshots and how our method addresses them. Decoding biochemical reaction networks–the number of species, reactions connecting them and reaction rates–from snapshot data, such as data drawn from single molecule fluorescence in situ hybridization (smFISH) experiments, is of critical interest. Yet, this task’s challenges are currently addressed with network specification heuristics since: 1) network size cannot be specified independently of rates; 2) rates may be separated by orders of magnitude, generating stiff ODEs; 3) uncertainty is not propagated into networks and parameters. We present, inspired by recent computational statistics tools, a method to simultaneously deduce reaction networks and associated rates from snapshot data while propagating error over all unknowns. We achieve this by treating network models as random variables and, within the Bayesian nonparametric paradigm, develop a posterior over models themselves. This multidimensional posterior naturally contains multiple hills and valleys, so we propose a combination of samplers allowing for the first simultaneous and self-consistent inference of networks and their associated rates. Our method’s ability to treat arbitrary numbers of states contrasts the current state of the art and may modify previous biological conclusions based on network-determination heuristics. We demonstrate on method to synthetic data mimicking smFISH experiments and demonstrate its improvement over naive MCMC schemes.
DOI: 10.1038/nbt.3269
发表时间: 2015-07-01
影响因子: 46.9
作者:
Gaidatzis, Dimos;Burger, Lukas;Stadler, Michael B.
通讯作者: Stadler, Michael B.
DOI: 10.1063/1.1472510
发表时间: 2002-05-22
影响因子: 4.4
作者:
Fukunishi, H;Watanabe, O;Takada, S
通讯作者: Takada, S
DOI: 10.1137/050628568
发表时间: 2006-01-01
影响因子: 3.1
作者:
Efendiev, Y.;Hou, T.;Luo, W.
通讯作者: Luo, W.
DOI: 10.1126/science.280.5363.585
发表时间: 1998-04-24
期刊: SCIENCE
影响因子: 56.9
作者:
Femino, A;Fay, FS;Singer, RH
通讯作者: Singer, RH
DOI: 10.1038/msb.2011.27
发表时间: 2011-05-24
影响因子: 9.9
作者:
通讯作者: --