A framework for evaluating epidemic forecasts.

A framework for evaluating epidemic forecasts.
复制标题

DOI:
10.1186/s12879-017-2365-1
复制
发表时间:
2017-05-15
影响因子:
3.7
通讯作者:
Marathe M
Marathe M
中科院分区:
医学3区
文献类型:
--
作者:
Tabataba FS;Chakraborty P;Ramakrishnan N;Venkatramanan S;Chen J;Lewis B;Marathe M

文献摘要

被引文献

相似文献

在过去的几十年里,在流行病预测领域提出了许多预测方法。这些方法可以分为不同的类别,如确定性方法与概率方法,比较方法与生成方法,等等。在一些比较流行的比较方法中,研究人员将疫情爆发早期阶段观察到的流行病学数据与预测该流行病未来趋势和流行程度的拟议模型的输出进行比较。该领域的一个重要问题是缺乏标准的、定义良好的评估措施来从不同的算法中选择最佳算法,以及为特定算法选择最佳可能配置。在本文中,我们提出了一个评估框架,该框架允许组合不同的特征,误差度量和排名模式来评估预测。我们描述了各种流行病特征(epi特征),包括表征预测方法的输出,并提供适当的误差测量,可用于评估方法相对于这些epi特征的准确性。我们侧重于长期预测而不是短期预测,并通过评估预测美国流感的六种预测方法来证明该框架的实用性。我们的结果表明,即使对于单个Epi-feature,不同的误差度量也会导致不同的排名。此外,我们的实验分析表明,在跨误差测量评估时,没有一种方法在预测所有epi特征方面占主导地位。作为替代方案,我们提供了各种汇总个人排名的共识排名模式,从而考虑到不同的误差度量。由于每个epi特征都反映了疫情的不同方面,因此需要结合多种方法来提供全面的预测。因此,我们呼吁在评估流行病预测时采用更细致入微的方法,我们相信,本文中提出的全面评估框架将为计算流行病学社区增加价值。本文的在线版本(doi:10.1186/s12879-017-2365-1)包含补充材料,授权用户可以使用。
Over the past few decades, numerous forecasting methods have been proposed in the field of epidemic forecasting. Such methods can be classified into different categories such as deterministic vs. probabilistic, comparative methods vs. generative methods, and so on. In some of the more popular comparative methods, researchers compare observed epidemiological data from the early stages of an outbreak with the output of proposed models to forecast the future trend and prevalence of the pandemic. A significant problem in this area is the lack of standard well-defined evaluation measures to select the best algorithm among different ones, as well as for selecting the best possible configuration for a particular algorithm. In this paper we present an evaluation framework which allows for combining different features, error measures, and ranking schema to evaluate forecasts. We describe the various epidemic features (Epi-features) included to characterize the output of forecasting methods and provide suitable error measures that could be used to evaluate the accuracy of the methods with respect to these Epi-features. We focus on long-term predictions rather than short-term forecasting and demonstrate the utility of the framework by evaluating six forecasting methods for predicting influenza in the United States. Our results demonstrate that different error measures lead to different rankings even for a single Epi-feature. Further, our experimental analyses show that no single method dominates the rest in predicting all Epi-features when evaluated across error measures. As an alternative, we provide various Consensus Ranking schema that summarize individual rankings, thus accounting for different error measures. Since each Epi-feature presents a different aspect of the epidemic, multiple methods need to be combined to provide a comprehensive forecast. Thus we call for a more nuanced approach while evaluating epidemic forecasts and we believe that a comprehensive evaluation framework, as presented in this paper, will add value to the computational epidemiology community. The online version of this article (doi:10.1186/s12879-017-2365-1) contains supplementary material, which is available to authorized users.