Augmenting Co-Training With Recommendations to Classify Human Rights Violations

Augmenting Co-Training With Recommendations to Classify Human Rights Violations
复制标题

DOI:
10.1109/bigdata47090.2019.9005478
复制
发表时间:
2019-12
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Ragini Kihlman;Maria Fasli
Ragini Kihlman;Maria Fasli
中科院分区:
其他
文献类型:
--
作者:
Ragini Kihlman;Maria Fasli

文献摘要

相似文献

最近,许多人权组织开始使用社交媒体来识别、收集和记录侵犯人权的行为。从庞大的社交网络数据库中手动提取相关数据是困难的,耗时且昂贵。此外,随着技术的发展,侵犯人权行为的背景和重要性已经并将随着时间的推移而发生变化,需要专家的建议才能对这些数据进行任何形式的定量分析。有一些应用程序和系统可以帮助将这些数据结构化到相关的类别中,但是检测潜在的潜在模式,找到类似的注释模式并不断升级系统以执行探索性分析需要很高的维护和成本。本文提出了一种解决方案,通过集成半监督学习(矩阵分解)和相似性度量算法来解决这个问题,将大型非结构化语料库分类为带有一种或多种类型的侵犯人权行为的故事。在过去的几十年里,推荐系统已经成为强大的机器学习工具,可以从数据中推断并提供增值内容。沿着相同的上下文,半监督算法减轻了存在相对小的标记训练数据但存在大的未标记数据集的情况。本文尝试将这两种算法联合收割机来发现未标记的受害者幸存者故事中的模式,并从其他类似的故事中推荐标签,从而更新初始标记集。使用最先进的评价指标的算法的效率进行评估。实验结果显示了新故事和标记故事之间的相关性。实验结果表明,该算法优于一些内部推荐算法。
In the recent past, many human rights organizations have started using social media to identify, collect and document human rights violations. To manually extract relevant data from the large corpus of this social network data is difficult and time-consuming and expensive. Furthermore, with the advent of technology, the context and significance of the human rights abuses has and will change over time and advice from experts is needed to perform any kind quantitative analysis on this data. There are applications and systems that help structure this data into relevant categories, but detecting underlying latent patterns, finding similar annotated patterns and continuously upgrading the system to perform exploratory analysis requires high maintenance and cost. This paper proposes a solution to address this problem by integrating semi-supervised learning (with Matrix Factorization) and similarity measures algorithms to classify the large unstructured corpus into stories that have been labelled with one or more types of human rights abuses. In the last few decades, recommender systems have come across as powerful machine learning tools to infer from data and provide value-added content. Along the same context, semi-supervised algorithms mitigate situations where there is a relatively small labelled training data, but a large unlabeled data-set. This paper tries to combine both these algorithms to discover patterns in unlabeled victim survivor stories and recommends labels from other similar stories, thus updating the initial labelled set. The efficiency of the algorithm is evaluated using state of art evaluation metrics. Experimental results show a correlation between new and labelled stories. Real-world results show that the algorithm outplays some of in house recommendation algorithms.