Deep learning for real-time social media text classification for situation awareness – using Hurricanes Sandy, Harvey, and Irma as case studies

Deep learning for real-time social media text classification for situation awareness – using Hurricanes Sandy, Harvey, and Irma as case studies
复制标题

DOI:
10.1080/17538947.2019.1574316
复制
发表时间:
2019-02
影响因子:
5.1
通讯作者:
Manzhu Yu;Qunying Huang;Han Qin;C. Scheele;C. Yang
Manzhu Yu;Qunying Huang;Han Qin;C. Scheele;C. Yang
中科院分区:
地球科学1区
文献类型:
--
作者:
Manzhu Yu;Qunying Huang;Han Qin;C. Scheele;C. Yang

文献摘要

被引文献

相似文献

摘要 在过去的几年中,社交媒体平台一直为灾害管理做出了贡献。使用传统机器学习技术的文本挖掘解决方案已被开发出来,将消息分类为不同的主题,例如警告和建议,以更好地理解含义并利用社交媒体文本内容中的有用信息。然而,这些方法大多是特定于事件的,很难推广到跨事件分类。换句话说,由历史数据集训练的传统分类模型无法对未来事件中的社交媒体消息进行分类。本研究基于飓风桑迪、哈维和艾尔玛期间收集的三个地理标记 Twitter 数据集,检验了卷积神经网络 (CNN) 模型在跨事件 Twitter 主题分类中的能力。将 CNN 模型的性能与两种传统机器学习方法进行了比较:支持向量机 (SVM) 和逻辑回归 (LR)。实验结果表明,CNN 模型在单事件和跨事件评估场景中均取得了较高的准确度,而 SVM 和 LR 模型与各自的单事件准确度结果相比,准确度较低。这表明 CNN 模型具有对过去事件中的 Twitter 数据进行预训练的能力,以便对即将发生的事件进行分类,以实现态势感知。
ABSTRACT Social media platforms have been contributing to disaster management during the past several years. Text mining solutions using traditional machine learning techniques have been developed to categorize the messages into different themes, such as caution and advice, to better understand the meaning and leverage useful information from the social media text content. However, these methods are mostly event specific and difficult to generalize for cross-event classifications. In other words, traditional classification models trained by historic datasets are not capable of categorizing social media messages from a future event. This research examines the capability of a convolutional neural network (CNN) model in cross-event Twitter topic classification based on three geo-tagged twitter datasets collected during Hurricanes Sandy, Harvey, and Irma. The performance of the CNN model is compared to two traditional machine learning methods: support vector machine (SVM) and logistic regression (LR). Experiment results showed that CNN models achieved a consistently better accuracy for both single event and cross-event evaluation scenarios whereas SVM and LR models had lower accuracy compared to their own single event accuracy results. This indicated that the CNN model has the capability of pre-training Twitter data from past events to classify for an upcoming event for situational awareness.