A multi-modal approach towards mining social media data during natural disasters - a case study of Hurricane Irma

A multi-modal approach towards mining social media data during natural disasters - a case study of Hurricane Irma
复制标题

DOI:
10.1016/j.ijdrr.2020.102032
复制
发表时间:
2021-01
期刊:
International journal of disaster risk reduction : IJDRR
影响因子:
--
通讯作者:
S. Mohanty;B. Biggers;S. SayedAhmed;Nastaran Pourebrahim;E. Goldstein;Rick L. Bunch;G. Chi;F. Sadri;Tom P. McCoy;A. Cosby
S. Mohanty;B. Biggers;S. SayedAhmed;Nastaran Pourebrahim;E. Goldstein;Rick L. Bunch;G. Chi;F. Sadri;Tom P. McCoy;A. Cosby
中科院分区:
其他
文献类型:
--
作者:
S. Mohanty;B. Biggers;S. SayedAhmed;Nastaran Pourebrahim;E. Goldstein;Rick L. Bunch;G. Chi;F. Sadri;Tom P. McCoy;A. Cosby

文献摘要

相似文献

流媒体社交媒体提供了极端天气影响的实时一瞥。然而,大量的流数据使挖掘信息成为应急管理人员、政策制定者和学科科学家的挑战。在这里,我们探讨了数据学习方法的有效性,以挖掘和过滤来自飓风厄玛在美国佛罗里达登陆的社交媒体数据流的信息。我们使用了来自16,598名用户的54,383条Twitter消息(来自784 K地理定位消息)。2017年10月10日至12日,开发4个独立的模型来过滤相关性数据:1)基于每个推文的地点和时间的强制条件的地理空间模型,2)包括图像的推文的图像分类模型,3)预测高音喇叭可靠性的用户模型,以及4)确定文本是否与飓风伊尔玛有关的文本模型。所有四个模型都经过独立测试,并且可以结合起来,根据用户为每个子模型定义的阈值快速过滤和可视化推文。我们设想,这种类型的过滤和可视化例程可以作为从噪声源(如Twitter)中捕获数据的基础模型。这些数据随后可以被政策制定者、环境管理者、应急管理者和领域科学家使用,他们有兴趣找到具有特定属性的推文,以便在灾难的不同阶段使用(例如,准备、响应和恢复)或详细研究。
Streaming social media provides a real-time glimpse of extreme weather impacts. However, the volume of streaming data makes mining information a challenge for emergency managers, policy makers, and disciplinary scientists. Here we explore the effectiveness of data learned approaches to mine and filter information from streaming social media data from Hurricane Irma's landfall in Florida, USA. We use 54,383 Twitter messages (out of 784 K geolocated messages) from 16,598 users from Sept. 10–12, 2017 to develop 4 independent models to filter data for relevance: 1) a geospatial model based on forcing conditions at the place and time of each tweet, 2) an image classification model for tweets that include images, 3) a user model to predict the reliability of the tweeter, and 4) a text model to determine if the text is related to Hurricane Irma. All four models are independently tested, and can be combined to quickly filter and visualize tweets based on user-defined thresholds for each submodel. We envision that this type of filtering and visualization routine can be useful as a base model for data capture from noisy sources such as Twitter. The data can then be subsequently used by policy makers, environmental managers, emergency managers, and domain scientists interested in finding tweets with specific attributes to use during different stages of the disaster (e.g., preparedness, response, and recovery), or for detailed research.