Augmenting Co-Training With Recommendations to Classify Human Rights Violations
Augmenting Co-Training With Recommendations to Classify Human Rights Violations
复制标题
DOI:
10.1109/bigdata47090.2019.9005478
复制
发表时间:
2019-12
期刊:
影响因子:
--
通讯作者:
Ragini Kihlman;Maria Fasli
中科院分区:
文献类型:
--
作者:
Ragini Kihlman;Maria Fasli
In the recent past, many human rights organizations have started using social media to identify, collect and document human rights violations. To manually extract relevant data from the large corpus of this social network data is difficult and time-consuming and expensive. Furthermore, with the advent of technology, the context and significance of the human rights abuses has and will change over time and advice from experts is needed to perform any kind quantitative analysis on this data. There are applications and systems that help structure this data into relevant categories, but detecting underlying latent patterns, finding similar annotated patterns and continuously upgrading the system to perform exploratory analysis requires high maintenance and cost. This paper proposes a solution to address this problem by integrating semi-supervised learning (with Matrix Factorization) and similarity measures algorithms to classify the large unstructured corpus into stories that have been labelled with one or more types of human rights abuses. In the last few decades, recommender systems have come across as powerful machine learning tools to infer from data and provide value-added content. Along the same context, semi-supervised algorithms mitigate situations where there is a relatively small labelled training data, but a large unlabeled data-set. This paper tries to combine both these algorithms to discover patterns in unlabeled victim survivor stories and recommends labels from other similar stories, thus updating the initial labelled set. The efficiency of the algorithm is evaluated using state of art evaluation metrics. Experimental results show a correlation between new and labelled stories. Real-world results show that the algorithm outplays some of in house recommendation algorithms.