Efficient Bayesian inference for stochastic agent-based models.

Efficient Bayesian inference for stochastic agent-based models.
复制标题

DOI:
10.1371/journal.pcbi.1009508
复制
发表时间:
2022-10
影响因子:
4.3
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

许多现实世界问题的建模依赖于对随机相互作用的个体或代理的大量计算模拟。然而,作为主体之间相互作用基础的参数值通常是鲜为人知的,因此它们需要从系统的宏观观察中推断出来。由于统计推断依赖于重复模拟来采样参数空间,因此这些模拟的高计算费用可能成为绊脚石。在本文中,我们比较了通过使用机器学习方法在贝叶斯环境中缓解这一问题的两种方法:一种方法是构建轻量级代理模型来替代推理中使用的模拟。或者,可以完全避开贝叶斯抽样方案的需要,直接估计后验分布。我们专注于跟踪自主代理的随机模拟,并提出两个案例研究:肿瘤生长和传染病传播。我们证明,通过相对少量的模拟可以实现良好的推理精度,使我们的机器学习方法比依赖于采样参数空间的基于模拟的经典方法快几个数量级。然而,我们发现,虽然一些方法通常比其他方法产生更稳健的结果,但当试图从观测中推断模型参数时,没有算法提供一个通用的解决方案。相反,必须考虑到具体的实际应用程序来选择推理技术。考虑到现实世界现象的随机性质,对一些方法来说,这是一个额外的挑战,可能变得无法克服。总的来说,我们发现创建直接推理机的机器学习方法在现实世界的应用中很有前途。我们将我们的研究结果作为建模从业人员的一般指导方针。计算机模拟在现代科学中起着至关重要的作用,因为它们通常用于将理论与观测进行比较。人们可以通过将数据与不同情况下的预测行为进行比较来推断系统的属性。每个场景对应一个设置略有不同的模拟。然而,由于现实世界的问题非常复杂,模拟通常需要大量的计算资源,这使得与数据的直接比较具有挑战性,如果不是无法克服的话。因此,有必要诉诸推理方法来缓解这一问题,但对于任何具体的研究问题,选择何种路径并不明确。在本文中,我们提供了如何做出这种选择的一般准则。我们通过研究肿瘤学和流行病学的例子以及利用机器学习来做到这一点。更具体地说,我们专注于跟踪自主代理(如单个细胞或个体)行为的模拟。我们展示了最好的前进方式是问题依赖的,并强调了在不同的案例研究中产生最可靠结果的方法。与其依赖单一的推理技术,我们建议采用几种方法,并根据预先确定的标准选择最可靠的方法。
The modelling of many real-world problems relies on computationally heavy simulations of randomly interacting individuals or agents. However, the values of the parameters that underlie the interactions between agents are typically poorly known, and hence they need to be inferred from macroscopic observations of the system. Since statistical inference rests on repeated simulations to sample the parameter space, the high computational expense of these simulations can become a stumbling block. In this paper, we compare two ways to mitigate this issue in a Bayesian setting through the use of machine learning methods: One approach is to construct lightweight surrogate models to substitute the simulations used in inference. Alternatively, one might altogether circumvent the need for Bayesian sampling schemes and directly estimate the posterior distribution. We focus on stochastic simulations that track autonomous agents and present two case studies: tumour growths and the spread of infectious diseases. We demonstrate that good accuracy in inference can be achieved with a relatively small number of simulations, making our machine learning approaches orders of magnitude faster than classical simulation-based methods that rely on sampling the parameter space. However, we find that while some methods generally produce more robust results than others, no algorithm offers a one-size-fits-all solution when attempting to infer model parameters from observations. Instead, one must choose the inference technique with the specific real-world application in mind. The stochastic nature of the considered real-world phenomena poses an additional challenge that can become insurmountable for some approaches. Overall, we find machine learning approaches that create direct inference machines to be promising for real-world applications. We present our findings as general guidelines for modelling practitioners. Computer simulations play a vital role in modern science as they are commonly used to compare theory with observations. One can infer the properties of a system by comparing the data to the predicted behaviour in different scenarios. Each scenario corresponds to a simulation with slightly different settings. However, since real-world problems are highly complex, the simulations often require extensive computational resources, making direct comparisons with data challenging, if not insurmountable. It is, therefore, necessary to resort to inference methods that mitigate this issue, but it is not clear-cut what path to choose for any specific research problem. In this paper, we provide general guidelines for how to make this choice. We do so by studying examples from oncology and epidemiology and by taking advantage of machine learning. More specifically, we focus on simulations that track the behaviour of autonomous agents, such as single cells or individuals. We show that the best way forward is problem-dependent and highlight the methods that yield the most robust results across the different case studies. Rather than relying on a single inference technique, we recommend employing several methods and selecting the most reliable based on predetermined criteria.
将机器学习作为基于代理的模拟的替代模型。
DOI: 10.1371/journal.pone.0263150
发表时间: 2022
期刊: PloS one
影响因子: 3.7
作者:
Angione C;Silverman E;Yaneske E
通讯作者: Yaneske E
DOI: 10.1111/j.1365-2966.2012.21818.x
发表时间: 2012-12-01
影响因子: 4.8
作者:
Bazot, M.;Bourguignon, S.;Christensen-Dalsgaard, J.
通讯作者: Christensen-Dalsgaard, J.
DOI: 10.1051/0004-6361/201015451
发表时间: 2011-03-01
影响因子: 6.5
作者:
Handberg, R.;Campante, T. L.
通讯作者: Campante, T. L.
DOI: 10.1038/s41591-020-1001-6
发表时间: 2020-07-14
期刊: NATURE MEDICINE
影响因子: 82.9
作者:
Hoertel, Nicolas;Blachier, Martin;Leleu, Henri
通讯作者: Leleu, Henri
DOI: 10.3847/0004-637x/830/1/31
发表时间: 2016-10-10
影响因子: 4.9
作者:
Bellinger, Earl P.;Angelou, George C.;Guggenberger, Elisabeth
通讯作者: Guggenberger, Elisabeth