Abusive Language Detection in Heterogeneous Contexts: Dataset Collection and the Role of Supervised Attention

Abusive Language Detection in Heterogeneous Contexts: Dataset Collection and the Role of Supervised Attention
复制标题

DOI:
10.1609/aaai.v35i17.17738
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Hongyu Gong;Alberto Valido;Katherine M. Ingram;G. Fanti;S. Bhat;D. Espelage
Hongyu Gong;Alberto Valido;Katherine M. Ingram;G. Fanti;S. Bhat;D. Espelage
中科院分区:
其他
文献类型:
--
作者:
Hongyu Gong;Alberto Valido;Katherine M. Ingram;G. Fanti;S. Bhat;D. Espelage

文献摘要

被引文献

相似文献

辱骂性语言是在线社交平台中的一个大问题。现有的辱骂性语言检测技术特别不适合于包含异质辱骂性语言模式的评论,即,包括虐待和非虐待的部分。这在一定程度上是由于缺乏数据集,明确地注释异质性的辱骂性语言。我们通过提供来自YouTube的11,000多条评论中的辱骂性语言的注释数据集来应对这一挑战。我们通过单独注释整个评论和组成每个评论的单个句子来解释这个数据集中的异质性。然后,我们提出了一种算法,使用监督注意力机制,使用多任务学习来检测和分类滥用内容。我们经验证明了使用传统的技术对异构内容的挑战和相对收益的性能所提出的方法超过国家的最先进的方法。
Abusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous abusive language patterns, i.e., both abusive and non-abusive parts. This is due in part to the lack of datasets that explicitly annotate heterogeneity in abusive language. We tackle this challenge by providing an annotated dataset of abusive language in over 11,000 comments from YouTube. We account for heterogeneity in this dataset by separately annotating both the comment as a whole and the individual sentences that comprise each comment. We then propose an algorithm that uses a supervised attention mechanism to detect and categorize abusive content using multi-task learning. We empirically demonstrate the challenges of using traditional techniques on heterogeneous content and the comparative gains in performance of the proposed approach over state-of-the-art methods.