Overview of the HASOC track at FIRE 2020: Hate Speech and Offensive Content Identification in Indo-European Languages

Overview of the HASOC track at FIRE 2020: Hate Speech and Offensive Content Identification in Indo-European Languages
复制标题

FIRE 2020 上的 HASOC 赛道概述:印欧语言中的仇恨言论和攻击性内容识别

DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Aditya Patel
Aditya Patel
中科院分区:
农林科学3区
文献类型:
--
作者:
Thomas Mandl;Sandip J Modha;Prasenjit Majumder;Daksh Patel;Mohana Dave;Chintak Mandalia;Aditya Patel

文献摘要

被引文献

相似文献

随着社交媒体的发展,仇恨言论的传播也在迅速增加。社交媒体在许多国家被广泛使用。仇恨言论也在这些国家蔓延。这带来了对多语言仇恨言论检测算法的需求。目前,这一领域的许多研究都是针对英语的。HASOC计划提供一个平台,开发和优化印地语、德语和英语的仇恨言论检测算法。数据集是从Twitter存档中收集的,并由机器学习系统预先分类。HASOC对所有三种语言都有两个子任务:任务A是一个二元分类问题(仇恨和不冒犯),而任务B是一个细粒度分类问题,分为三类(仇恨)仇恨言论,冒犯和亵渎。总共有40支队伍参加了252次比赛。任务A的最佳分类算法的性能为F1,分别为英语、印地语和德语的0.51、0.53和0.52。对于任务B,对于英语、印地语和德语,最佳分类算法分别获得了0.26、0.33和0.29的F1测度。本文介绍了任务和数据开发以及结果。性能最好的算法主要是变压器结构BERT的变体。然而,其他系统也得到了很好的应用
With the growth of social media, the spread of hate speech is also increasing rapidly. Social media are widely used in many countries. Also Hate Speech is spreading in these countries. This brings a need for multilingual Hate Speech detection algorithms. Much research in this area is dedicated to English at the moment. The HASOC track intends to provide a platform to develop and optimize Hate Speech detection algorithms for Hindi, German and English. The dataset is collected from a Twitter archive and pre-classified by a machine learning system. HASOC has two sub-task for all three languages: task A is a binary classification problem (Hate and Not Offensive) while task B is a fine-grained classification problem for three classes (HATE) Hate speech, OFFENSIVE and PROFANITY. Overall, 252 runs were submitted by 40 teams. The performance of the best classification algorithms for task A are F1 measures of 0.51, 0.53 and 0.52 for English, Hindi, and German, respectively. For task B, the best classification algorithms achieved F1 measures of 0.26, 0.33 and 0.29 for English, Hindi, and German, respectively. This article presents the tasks and the data development as well as the results. The best performing algorithms were mainly variants of the transformer architecture BERT. However, also other systems were applied with good success