Bayesian history matching of complex infectious disease models using emulation: a tutorial and a case study on HIV in Uganda.

Bayesian history matching of complex infectious disease models using emulation: a tutorial and a case study on HIV in Uganda.
复制标题

DOI:
10.1371/journal.pcbi.1003968
复制
发表时间:
2015-01
影响因子:
4.3
通讯作者:
White RG
White RG
中科院分区:
生物学2区
文献类型:
--
作者:
Andrianakis I;Vernon IR;McCreesh N;McKinley TJ;Oakley JE;Nsubuga RN;Goldstein M;White RG

文献摘要

参考文献

被引文献

相似文献

科学计算的进步使得复杂模型得以发展,这些模型经常被应用于疾病流行病学、公共卫生和决策制定等问题。这些模型的实用性部分取决于它们对经验数据的再现能力。然而,大量的输入和输出参数以及较长的运行时间极大地阻碍了将此类模型与现实世界的数据进行拟合,以至于许多建模研究缺乏正式的校准方法。我们提出了一种新方法,它有可能改进复杂传染病模型(以下称为模拟器)的校准。我们以教程和案例研究的形式呈现这一方法,在案例研究中,我们利用乌干达现有的大量人口、行为和流行病学数据,对一个动态的、事件驱动的、基于个体的随机艾滋病病毒模拟器进行历史匹配。该教程描述了历史匹配和仿真。历史匹配是一种迭代程序,它通过识别和舍弃不太可能与经验数据良好匹配的区域来缩小模拟器的输入空间。历史匹配依赖于模拟器的贝叶斯表示形式(称为仿真器)的计算效率。仿真器模拟模拟器的行为,但评估速度通常要快几个数量级。在案例研究中,我们使用一个有22个输入的模拟器,同时拟合其18个输出。经过9次历史匹配迭代,确定了模拟器输入空间的一个合理区域,该区域比原始输入空间小[此处缺失倍数相关内容]倍。发现在该区域内进行的模拟器评估有65%的概率拟合所有18个输出。历史匹配和仿真是传染病建模者工具包中有用的补充。还需要进一步研究以明确解决模拟器的随机性以及考虑输出之间的相关性。 越来越多的学科,特别是生物学,依赖复杂的计算模型。这些模型的实用性取决于它们与经验数据的拟合程度。拟合是通过为模型的输入参数寻找合适的值来实现的,这一过程称为校准。现代计算机模型通常有大量的输入和输出参数以及较长的运行时间,这是其计算复杂性不断增加的结果。上述两点阻碍了校准过程。在这项工作中,我们提出了一种方法,可以帮助校准具有较长运行时间以及多个输入和输出的模型。我们将这种方法应用于一个基于个体的、动态的和随机的艾滋病病毒模型,使用来自乌干达的艾滋病病毒数据。最终系统有65%的概率选择一个能拟合所有18个模型输出的输入参数集。
Advances in scientific computing have allowed the development of complex models that are being routinely applied to problems in disease epidemiology, public health and decision making. The utility of these models depends in part on how well they can reproduce empirical data. However, fitting such models to real world data is greatly hindered both by large numbers of input and output parameters, and by long run times, such that many modelling studies lack a formal calibration methodology. We present a novel method that has the potential to improve the calibration of complex infectious disease models (hereafter called simulators). We present this in the form of a tutorial and a case study where we history match a dynamic, event-driven, individual-based stochastic HIV simulator, using extensive demographic, behavioural and epidemiological data available from Uganda. The tutorial describes history matching and emulation. History matching is an iterative procedure that reduces the simulator's input space by identifying and discarding areas that are unlikely to provide a good match to the empirical data. History matching relies on the computational efficiency of a Bayesian representation of the simulator, known as an emulator. Emulators mimic the simulator's behaviour, but are often several orders of magnitude faster to evaluate. In the case study, we use a 22 input simulator, fitting its 18 outputs simultaneously. After 9 iterations of history matching, a non-implausible region of the simulator input space was identified that was times smaller than the original input space. Simulator evaluations made within this region were found to have a 65% probability of fitting all 18 outputs. History matching and emulation are useful additions to the toolbox of infectious disease modellers. Further research is required to explicitly address the stochastic nature of the simulator as well as to account for correlations between outputs. An increasing number of scientific disciplines, and biology in particular, rely on complex computational models. The utility of these models depends on how well they are fitted to empirical data. Fitting is achieved by searching for suitable values for the models' input parameters, in a process known as calibration. Modern computer models typically have a large number of input and output parameters, and long running times, a consequence of their increasing computational complexity. The above two things hinder the calibration process. In this work, we propose a method that can help the calibration of models with long running times and several inputs and outputs. We apply this method on an individual based, dynamic and stochastic HIV model, using HIV data from Uganda. The final system has a 65% probability of selecting an input parameter set that fits all 18 model outputs.
DOI: 10.1098/rsif.2007.1292
发表时间: 2008-08-06
影响因子: 3.9
作者:
Cauchemez, Simon;Ferguson, Neil M.
通讯作者: Ferguson, Neil M.
DOI: 10.1016/j.csda.2012.04.020
发表时间: 2012-12-01
影响因子: 1.8
作者:
Andrianakis, Ioannis;Challenor, Peter G.
通讯作者: Challenor, Peter G.
DOI: 10.1016/j.jspi.2009.08.006
发表时间: 2010-03-01
影响因子: 0.9
作者:
Conti, Stefano;O'Hagan, Anthony
通讯作者: O'Hagan, Anthony
DOI: 10.1198/jasa.2009.0005
发表时间: 2009-03-01
影响因子: 3.7
作者:
Henderson, Daniel A.;Boys, Richard J.;Wilkinson, Darren J.
通讯作者: Wilkinson, Darren J.
DOI: 10.1111/j.1365-2966.2010.16991.x
发表时间: 2010-10-01
影响因子: 4.8
作者:
Bower, R. G.;Vernon, I.;Frenk, C. S.
通讯作者: Frenk, C. S.