Identifying valuable information from twitter during natural disasters

Identifying valuable information from twitter during natural disasters
复制标题

在自然灾害期间从 Twitter 中识别有价值的信息

DOI:
--
复制
发表时间:
2014
期刊:
ASIS&T Annual Meeting
影响因子:
--
通讯作者:
Andrea H. Tapia
Andrea H. Tapia
中科院分区:
--
文献类型:
--
作者:
Brandon Truong;Cornelia Caragea;A. Squicciarini;Andrea H. Tapia

文献摘要

被引文献

相似文献

在任何重大事件中,社交媒体都是重要的信息来源,尤其是自然灾害。然而,随着社交媒体数据量的指数级增长,不能提供有价值信息的会话数据也随之增加,特别是在灾难事件的背景下,因此,人们找到组织救援工作、寻求帮助和潜在拯救生命所需信息的能力减弱。这个项目的重点是开发一种贝叶斯方法来对飓风桑迪期间的推文(Twitter上的帖子)进行分类,以区分“信息性”和“会话性”推文。我们设计了一组有效的特征,并将它们作为朴素贝叶斯分类器的输入。与“单词包”方法相比,新功能集在tweet分类方面提供了类似的结果。然而,设计的特征集只包含9个特征,而“词包”有3000多个特征。当特征集与“词袋”相结合时,准确率达到85.2914%。如果将我们的方法集成到与灾害相关的系统中,我们的方法可以为任何在自然灾害中寻求提取有用信息的个人或组织带来福音。
Social media is a vital source of information during any major event, especially natural disasters. However, with the exponential increase in volume of social media data, so comes the increase in conversational data that does not provide valuable information, especially in the context of disaster events, thus, diminishing peoples’ ability to find the information that they need in order to organize relief efforts, find help, and potentially save lives. This project focuses on the development of a Bayesian approach to the classification of tweets (posts on Twitter) during Hurricane Sandy in order to distinguish “informational” from “conversational” tweets. We designed an effective set of features and used them as input to Naive Bayes classifiers. In comparison to a “bag of words” approach, the new feature set provides similar results in the classification of tweets. However, the designed feature set contains only 9 features compared with more than 3000 features for “bag of words.” When the feature set is combined with “bag of words”, accuracy achieves 85.2914%. If integrated into disaster-related systems, our approach can serve as a boon to any person or organization seeking to extract useful information in the midst of a natural disaster.