Tweets Classification with BERT in the Field of Disaster Management

Tweets Classification with BERT in the Field of Disaster Management
复制标题

灾害管理领域使用 BERT 进行推文分类

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Guoqin Ma
Guoqin Ma
中科院分区:
--
文献类型:
--
作者:
Guoqin Ma

文献摘要

被引文献

相似文献

危机信息学关注用户生成内容(UGC)对灾害管理的贡献。为了有效地利用社交媒体数据,从大量数据流中过滤出噪声信息至关重要,以便我们可以更好地估计这些数据的灾害损失。不满足于基本的基于关键字的过滤,许多研究人员转向机器学习来解决问题。在这个项目中,我应用深度学习技术来解决灾害管理领域的推文分类问题。推特的标签反映了不同类型的灾害相关信息,这些信息在应急响应中具有不同的潜在用途。BERT用于迁移学习。用于分类的标准BERT架构和其他几个定制的BERT架构经过训练,以与具有预训练的Glove Twitter嵌入的基线双向LSTM进行比较。结果表明,BERT和基于BERT的LSTM获得了最好的结果,在F-1得分方面平均分别超过基线模型3.29%。模糊性和主观性对这些模型的性能有很大影响。在一些例子中,模型可以超越人类的表现。
Crisis informatics focus on the contribution of user generated content (UGC) to disaster management. To leverage the social media data effectively, it is crucial to filter out noisy information from the large volume of data flow so that we could better estimate disaster damage with these data. Not satisfied with basic keyword-based filtration, many researchers turn to machine learning for solution. In this project, I apply deep learning techniques to address Tweets classification problem in disaster management field. The labels of Tweets reflect different types of disaster-related information, which have different potential usage in emergency response. In particular, BERT is used for transfer learning. The standard BERT architecture for classification and several other customized BERT architectures are trained to compare with the baseline bidirectional LSTM with pretrained Glove Twitter embeddings. Results show that BERT and BERT-based LSTM attain the best results, outperforming the baseline model by 3.29% on average in terms of F-1 score respectively. Ambiguity and subjectivity affect the performance of these models considerably. In some examples the models can surpass human performance.