Enabling Rapid Classification of Social Media Communications During Crises

Enabling Rapid Classification of Social Media Communications During Crises
复制标题

在危机期间实现社交媒体通信的快速分类

DOI:
--
复制
发表时间:
2016
期刊:
International Journal of Information Systems for Crisis Response and Management
影响因子:
--
通讯作者:
Jaideep Srivastava
Jaideep Srivastava
中科院分区:
--
文献类型:
--
作者:
Muhammad Imran;P. Mitra;Jaideep Srivastava

文献摘要

参考文献

被引文献

相似文献

危机期间受影响人群对 Twitter 等社交媒体平台的使用被认为是应对危机的重要信息来源。然而,快速的危机应对需要对在线信息进行实时分析。当灾难发生时,除了其他数据处理技术之外,监督机器学习可以帮助实时对在线信息进行分类。然而,标记数据的稀缺导致机器训练的性能不佳。通常可以使用过去事件的标记数据。过去的标记数据可以重新用于训练分类器吗?我们研究过去事件的标记数据的有用性。我们观察使用从过去灾难中获得的训练集的不同组合来训练的分类器的性能。此外,我们提出了两种方法(目标标记和主动学习)来提高学习方案的分类性能。我们对真实危机数据集进行了广泛的实验,并展示了过去标记数据在训练机器学习分类器以实时处理突发危机相关数据方面的实用性。
The use of social media platforms such as Twitter by affected people during crises is considered a vital source of information for crisis response. However, rapid crisis response requires real-time analysis of online information. When a disaster happens, among other data processing techniques, supervised machine learning can help classify online information in real-time. However, scarcity of labeled data causes poor performance in machine training. Often labeled data from past event is available. Can past labeled data be reused to train classifiers? We study the usefulness of labeled data of past events. We observe the performance of our classifiers trained using different combinations of training sets obtained from past disasters. Moreover, we propose two approaches (target labeling and active learning) to boost classification performance of a learning scheme. We perform extensive experimentation on real crisis datasets and show the utility of past-labeled data to train machine learning classifiers to process sudden-onset crisis-related data in real-time.
DOI: 10.1056/nejmp0900702
发表时间: 2009-05-21
期刊: The New England journal of medicine
影响因子: --
作者:
Brownstein JS;Freifeld CC;Madoff LC
通讯作者: Madoff LC