Classification of Twitter Disaster Data Using a Hybrid Feature-Instance Adaptation Approach

Classification of Twitter Disaster Data Using a Hybrid Feature-Instance Adaptation Approach
复制标题

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
R. Mazloom;Hongmin Li;Doina Caragea;Muhammad Imran;Cornelia Caragea
R. Mazloom;Hongmin Li;Doina Caragea;Muhammad Imran;Cornelia Caragea
中科院分区:
其他
文献类型:
--
作者:
R. Mazloom;Hongmin Li;Doina Caragea;Muhammad Imran;Cornelia Caragea

文献摘要

相似文献

社交媒体在紧急情况下产生的大量数据被视为关键信息的宝库。由于缺乏特定灾难的标签数据,在灾难的早期阶段使用有监督的机器学习技术受到了挑战。此外,考虑到当前灾难和先前灾难之间的内在差异,基于先前灾难的标记数据训练的监督模型可能不会产生准确的结果。为了应对目标灾难缺乏标签数据带来的挑战,我们提出了一种分别基于矩阵分解和k近邻算法的混合特征-实例自适应方法。提出的混合适应方法被用来选择代表目标灾难的源灾难数据的子集。所选择的子集随后被用于学习目标灾难的准确的朴素贝叶斯分类器。
Huge amounts of data that are generated on social media during emergency situations are regarded as troves of critical information. The use of supervised machine learning techniques in the early stages of a disaster is challenged by the lack of labeled data for that particular disaster. Furthermore, supervised models trained on labeled data from a prior disaster may not produce accurate results, given the inherent variation between the current and the prior disasters. To address the challenges posed by the lack of labeled data for a target disaster, we propose to use a hybrid feature-instance adaptation approach based on matrix factorization and the k-nearest neighbors algorithm, respectively. The proposed hybrid adaptation approach is used to select a subset of the source disaster data that is representative for the target disaster. The selected subset is subsequently used to learn accurate Naïve Bayes classifiers for the target disaster.