Offensive Comments in the Brazilian Web: a dataset and baseline results

Offensive Comments in the Brazilian Web: a dataset and baseline results
复制标题

巴西网络中的攻击性评论:数据集和基线结果

DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
V. Moreira
V. Moreira
中科院分区:
--
文献类型:
--
作者:
Rogers Pelle;V. Moreira

文献摘要

被引文献

相似文献

.巴西的网络用户是社交网络中最活跃的群体之一,他们非常热衷于与他人互动。被称为仇恨言论的攻击性言论一直困扰着在线媒体,并引发了一系列针对发布网络内容的公司的法律诉讼。鉴于每天发布的大量用户生成的文本,手动过滤攻击性评论变得不可行。攻击性评论的识别可以被视为一项监督分类任务。为了获得对评论进行分类的模型,需要包含正面和负面示例的注释数据集。在葡萄牙语中缺乏这样的数据集,限制了这种语言的检测方法的发展。在本文中,我们描述了我们如何创建注释数据集的攻击性评论葡萄牙语收集巴西网络上的新闻评论。此外,我们还提供了通过标准分类算法在这些数据集上实现的分类结果,这些结果可以作为该主题未来工作的基线。
. Brazilian Web users are among the most active in social networks and very keen on interacting with others. Offensive comments, known as hate speech , have been plaguing online media and originating a number of law-suits against companies which publish Web content. Given the massive number of user generated text published on a daily basis, manually filtering offensive comments becomes infeasible. The identification of offensive comments can be treated as a supervised classification task. In order to obtain a model to classify comments, an annotated dataset containing positive and negative examples is necessary. The lack of such a dataset in Portuguese, limits the development of detection approaches for this language. In this paper, we describe how we created annotated datasets of offensive comments for Portuguese by collecting news comments on the Brazilian Web. In addition, we provide classification re-sults achieved by standard classification algorithms on these datasets which can serve as baseline for future work on this topic.