A Comparative Study of Deep Learning Methods for Hate Speech and Offensive Language Detection in Textual Data

A Comparative Study of Deep Learning Methods for Hate Speech and Offensive Language Detection in Textual Data
复制标题

文本数据中仇恨言论和攻击性语言检测的深度学习方法的比较研究

DOI:
10.1109/indicon52576.2021.9691704
复制
发表时间:
2021
期刊:
2021 IEEE 18th India Council International Conference (INDICON)
影响因子:
--
通讯作者:
R. Sinha
R. Sinha
中科院分区:
--
文献类型:
--
作者:
Yogesh Yadav;Parth Bajaj;Rohan Kumar Gupta;R. Sinha

文献摘要

被引文献

相似文献

社交网站上的仇恨言论问题非常普遍,每个主要社交媒体平台都面临着这一问题。已经探索了几种方法用于基于意图的文本分类的目的。每种方法都有自己的优点和缺点的类型的意图,数据集的大小,文本的最大长度等几种方法已经在文献中提出的仇恨和攻击性的语音检测。这项工作的主要目标是对用于仇恨言论和攻击性语言检测的精选深度学习方法进行比较研究。这些方法包括递归神经网络(RNN)、卷积神经网络(CNN)、长短期记忆(LSTM)和来自Transformer(BERT)的双向编码器表示。我们研究了类加权技术对深度学习方法性能的影响。我们的研究发现,预训练的BERT模型在未加权和加权仇恨语音分类的情况下都优于其他探索的模型。对于攻击性语言分类,RNN和CNN模型分别在未加权和加权的情况下优于所有其他模型。结果表明,类加权技术大大提高了所有四种模型对仇恨言论的分类性能。
The problem of hate speech on social network sites is very prevalent which is being faced by every major social media platform. Several methods have been explored for the purpose of intent-based text classification. Each method has its own pros and cons concerning the type of intent, size of data set, the maximum length of text, etc. Several approaches have been presented in the literature for the hate and offensive speech detection. The main objective of this work is to present a comparative study among select deep learning methods for hate speech and offensive language detection. These methods include recurrent neural network (RNN), convolutional neural network (CNN), long shortterm memory (LSTM) and bidirectional encoder representations from transformer (BERT). We have investigated the effect of class weighting technique on the performance of the deep learning methods. Our study finds that the pre-trained BERT model outperforms the other explored models in case of both unweighted and weighted hate speech classification. For offensive language classification, RNN and CNN model outperforms all other models in case of unweighted and weighted respectively. It came out that, the class weighting technique has considerably boost the classification performance of all four models for hate speech.