Predicting gasoline shortage during disasters using social media

Predicting gasoline shortage during disasters using social media
复制标题

DOI:
10.1007/s00291-019-00559-8
复制
发表时间:
2020-09-01
期刊:
影响因子:
2.7
通讯作者:
Batta, Rajan
Batta, Rajan
中科院分区:
管理学4区
文献类型:
--
作者:
Khare, Abhinav;He, Qing;Batta, Rajan

文献摘要

被引文献

相似文献

在飓风等预测灾害发生时,汽油短缺是一种常见现象。对未来汽油短缺的预测可以指导各机构将供应推到正确的地区并缓解短缺。我们展示了如何将社交媒体数据纳入汽油供应决策。我们开发了一种系统的方法来检查社交媒体帖子,如推文,并感知未来的汽油短缺。我们建立了一个四阶段短缺预测方法。在第一阶段,我们过滤掉与汽油相关的推文。在第二阶段中,我们使用基于SVM的推文分类器来分类关于汽油短缺的推文,使用主题建模技术识别的一元词和主题作为我们的特征。在第三阶段中,我们预测的数量,未来推关于汽油短缺的混合损失函数,这是建立联合收割机ARIMA和泊松回归方法相结合。在第四阶段,我们采用泊松回归预测短缺使用的推文预测的数量在第三阶段。为了验证该方法,我们开发了一个案例研究,预测汽油短缺,使用在佛罗里达的发病和飓风厄玛登陆后产生的推文。我们将预测与Irma期间汽油短缺的地面事实进行了比较,根据常用的误差估计,结果非常准确。
Shortage of gasoline is a common phenomenon during onset of forecasted disasters like hurricanes. Prediction of future gasoline shortage can guide agencies in pushing supplies to the correct regions and mitigating the shortage. We demonstrate how to incorporate social media data into gasoline supply decision making. We develop a systematic approach to examine social media posts like tweets and sense future gasoline shortage. We build a four-stage shortage prediction methodology. In the first stage, we filter out tweets related to gasoline. In the second stage, we use an SVM-based tweet classifier to classify tweets about the gasoline shortage, using unigrams and topics identified using topic modeling techniques as our features. In the third stage, we predict the number of future tweets about gasoline shortage using a hybrid loss function, which is built to combine ARIMA and Poisson regression methods. In the fourth stage, we employ Poisson regression to predict shortage using the number of tweets predicted in the third stage. To validate the methodology, we develop a case study that predicts the shortage of gasoline, using tweets generated in Florida during the onset and post landfall of Hurricane Irma. We compare the predictions to the ground truth about gasoline shortage during Irma, and the results are very accurate based on commonly used error estimates.