Assessing the performance of real-time epidemic forecasts: A case study of Ebola in the Western Area region of Sierra Leone, 2014-15

Assessing the performance of real-time epidemic forecasts: A case study of Ebola in the Western Area region of Sierra Leone, 2014-15
复制标题

DOI:
10.1371/journal.pcbi.1006785
复制
发表时间:
2019-02-01
影响因子:
4.3
通讯作者:
Edmunds, W. John
Edmunds, W. John
中科院分区:
生物学2区
文献类型:
--
作者:
Funk, Sebastian;Camacho, Anton;Edmunds, W. John

文献摘要

被引文献

相似文献

基于数学模型的实时预测可以为传染病爆发期间的关键决策提供信息。然而,流行病预测很少在活动期间或活动之后进行评估,而且几乎没有关于评估最佳指标的指导。在这里,我们提出了一种评估方法,解开预测能力的不同组成部分,分别评估预测的校准,清晰度和偏见的指标。这不仅可以评估预测与现实的接近程度,还可以评估不确定性的量化程度。我们使用这种方法来分析我们在2013-16年西非埃博拉疫情期间为塞拉利昂西部地区生成的真实的每周预测的表现。我们研究了一系列基于当时半机械模型生成的模型拟合的预测模型变体,发现在未来一两周的短时间范围内可以实现良好的概率校准,但模型预测在较长的预测范围内越来越不可靠。这表明,预测的质量可能已经足够好,可以根据提前几周但不是更长时间的预测为决策提供信息,反映了推动流行病发展轨迹的过程中的高度不确定性。将基于半机械模型的预测与基于更简单的零模型的预测进行比较表明,最佳半机械模型变体在概率校准方面的表现优于零模型,并且这将在爆发的最早阶段就被确定。随着预测成为公共卫生工具包的一个常规部分,绩效评估标准对于评估数学模型的质量和提高其可信度以及在旨在做出最有用和最可靠的预测时阐明困难和权衡将非常重要。
Real-time forecasts based on mathematical models can inform critical decision-making during infectious disease outbreaks. Yet, epidemic forecasts are rarely evaluated during or after the event, and there is little guidance on the best metrics for assessment. Here, we propose an evaluation approach that disentangles different components of forecasting ability using metrics that separately assess the calibration, sharpness and bias of forecasts. This makes it possible to assess not just how close a forecast was to reality but also how well uncertainty has been quantified. We used this approach to analyse the performance of weekly forecasts we generated in real time for Western Area, Sierra Leone, during the 2013-16 Ebola epidemic in West Africa. We investigated a range of forecast model variants based on the model fits generated at the time with a semi-mechanistic model, and found that good probabilistic calibration was achievable at short time horizons of one or two weeks ahead but model predictions were increasingly unreliable at longer forecasting horizons. This suggests that forecasts may have been of good enough quality to inform decision making based on predictions a few weeks ahead of time but not longer, reflecting the high level of uncertainty in the processes driving the trajectory of the epidemic. Comparing forecasts based on the semi-mechanistic model to simpler null models showed that the best semi-mechanistic model variant performed better than the null models with respect to probabilistic calibration, and that this would have been identified from the earliest stages of the outbreak. As forecasts become a routine part of the toolkit in public health, standards for evaluation of performance will be important for assessing quality and improving credibility of mathematical models, and for elucidating difficulties and trade-offs when aiming to make the most useful and reliable forecasts.