Combining Search, Social Media, and Traditional Data Sources to Improve Influenza Surveillance.

Combining Search, Social Media, and Traditional Data Sources to Improve Influenza Surveillance.
复制标题

DOI:
10.1371/journal.pcbi.1004513
复制
发表时间:
2015-10
影响因子:
4.3
通讯作者:
Brownstein JS
Brownstein JS
中科院分区:
生物学2区
文献类型:
--
作者:
Santillana M;Nguyen AT;Dredze M;Paul MJ;Nsoesie EO;Brownstein JS

文献摘要

被引文献

相似文献

我们提出了一种基于机器学习的方法,能够通过利用来自多个数据源的数据来提供美国流感活动的实时(“nowcast”)和预测估计,这些数据源包括:谷歌搜索,Twitter微博,近实时的医院访问记录,以及来自参与式监测系统的数据。我们的主要贡献包括将多个流感样疾病(ILI)活动估计值(由每个数据源独立生成)结合起来,利用机器学习集成方法对ILI进行单一预测。我们的方法利用每个数据源中的信息,并在CDC的ILI报告发布之前长达四周的时间内产生准确的每周ILI预测。我们评估了我们的集成方法在2013-2014年(回顾)和2014-2015年(现场)流感季节的预测能力,每周四次。我们的集成方法展示了几个优点:(1)我们的集成方法的预测优于每个独立使用每个数据源的预测,(2)我们的方法可以提前一周产生预测GFT的实时估计具有可比的准确性,(3)我们的两周和三周预测估计具有可比的准确性使用自回归模型的实时预测。此外,我们的研究结果表明,通过将不同的数据流(以社交媒体和众包数据的形式)纳入所有时间范围内的流感预测,可以获得相当多的洞察力。互联网用户的聚合活动模式使得能够检测和跟踪多个人口范围的事件,例如疾病爆发、金融市场表现和在线电影选择中的偏好。因此,在过去十年中,提出了一系列旨在实时监测和预测这些事件的数学模型。随着我们发现适合跟踪这些事件的新方法和数据源,尚不清楚更多的信息是否会导致预测的改进。在人群水平的数字疾病检测的背景下,我们表明,将来自美国多个流感活动预测因子的信息联合收割机结合起来,而不是简单地选择表现最好的流感预测因子,是有利的。我们的研究结果表明,来自多个数据源的信息,如谷歌搜索,Twitter微博,几乎实时的医院访问记录,以及来自参与式监测系统的数据,相互补充,并在最佳组合时产生最准确和最强大的流感预测集。
We present a machine learning-based methodology capable of providing real-time (“nowcast”) and forecast estimates of influenza activity in the US by leveraging data from multiple data sources including: Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system. Our main contribution consists of combining multiple influenza-like illnesses (ILI) activity estimates, generated independently with each data source, into a single prediction of ILI utilizing machine learning ensemble approaches. Our methodology exploits the information in each data source and produces accurate weekly ILI predictions for up to four weeks ahead of the release of CDC’s ILI reports. We evaluate the predictive ability of our ensemble approach during the 2013–2014 (retrospective) and 2014–2015 (live) flu seasons for each of the four weekly time horizons. Our ensemble approach demonstrates several advantages: (1) our ensemble method’s predictions outperform every prediction using each data source independently, (2) our methodology can produce predictions one week ahead of GFT’s real-time estimates with comparable accuracy, and (3) our two and three week forecast estimates have comparable accuracy to real-time predictions using an autoregressive model. Moreover, our results show that considerable insight is gained from incorporating disparate data streams, in the form of social media and crowd sourced data, into influenza predictions in all time horizons. The aggregated activity patterns of Internet users have enabled the detection and tracking of multiple population-wide events such as disease outbreaks, financial markets performance, and preferences in online movie selections. As a consequence, a collection of mathematical models aiming at monitoring and predicting these events in real-time have been proposed in the past decade. As we discover new methods and data sources suitable to track these events, it is not clear whether more information will lead to improved predictions. In the context of digital disease detection at the population level, we show that it is advantageous to combine the information from multiple flu activity predictors in the US instead of simply choosing the best performing flu predictor. Our findings suggest that the information from multiple data sources such as Google searches, Twitter microblogs, nearly real-time hospital visit records, and data from a participatory surveillance system, complement one another and produce the most accurate and robust set of flu predictions when combined optimally.