Machine-learning methods for identifying social media-based requests for urgent help during hurricanes

Machine-learning methods for identifying social media-based requests for urgent help during hurricanes
复制标题

DOI:
10.1016/j.ijdrr.2020.101757
复制
发表时间:
2020-12-01
影响因子:
5
通讯作者:
Dontula, Aman
Dontula, Aman
中科院分区:
地球科学2区
文献类型:
--
作者:
Devaraj, Ashwin;Murthy, Dhiraj;Dontula, Aman

文献摘要

被引文献

相似文献

在大规模自然灾害期间,人们越来越多地使用社交媒体请求紧急帮助。以前的工作已经成功地应用机器学习分类器来检测粗粒度类别的推文,例如灾害类型和相关性。然而,缺乏专注于检测包含第一响应者可操作的帮助请求的推文的工作。使用2017年美国休斯顿飓风哈维期间发布的超过500万条推文,我们表明,虽然这样的请求是不常见的,他们往往生死攸关的性质证明了鸣叫分类器的发展,以检测他们。我们发现,性能最好的分类器是在词嵌入上训练的卷积神经网络(CNN),在平均词嵌入上训练的支持向量机(SVM),以及在unigrams和词性(POS)标签组合上训练的多层感知器(MLP)。这些模型的F1得分超过0.86,证实了它们在检测紧急推文方面的有效性。我们强调了平均词嵌入在训练非神经模型中的实用性,并且这些特征产生的结果与更传统的n-gram和POS特征具有竞争力。
Social media is increasingly used by people during large-scale natural disasters to request emergency help. Previous work has had success in applying machine-learning classifiers to detect tweets in coarse-grained categories, such as disaster type and relevance. However, there is a dearth of work that focuses on detecting tweets containing requests for help that are actionable by first responders. Using over 5 million tweets posted during 2017's Hurricane Harvey in Houston, U.S., we show that though such requests are uncommon, their often life-or death nature justifies the development of tweet classifiers to detect them. We find that the best-performing classifiers are a convolutional neural network (CNN) trained on word embeddings, support vector machine (SVM) trained on average word embeddings, and multilayer perceptron (MLP) trained on a combination of unigrams and part-of-speech (POS) tags. These models achieve F1 scores of over 0.86, confirming their efficacy in detecting urgent tweets. We highlight the utility of average word embeddings for training non-neural models, and that such features produce results competitive with more traditional n-gram and POS features.