Hate Speech detection in English and Malayalam Code-Mixed Text using BERT embedding

Hate Speech detection in English and Malayalam Code-Mixed Text using BERT embedding
复制标题

使用 BERT 嵌入检测英语和马拉雅拉姆语代码混合文本中的仇恨言论

DOI:
10.1109/ic3sis54991.2022.9885339
复制
发表时间:
2022
期刊:
2022 International Conference on Computing, Communication, Security and Intelligent Systems (IC3SIS)
影响因子:
--
通讯作者:
Akhil Madhu
Akhil Madhu
中科院分区:
--
文献类型:
--
作者:
P. Deepasree Varma;P. Vinod;M. Nandakumar;K. Akshay;Akhil Madhu

文献摘要

被引文献

相似文献

仇恨语音检测是近几年来非常热门的研究领域。不同的研究者给仇恨言论下了不同的定义。在本文中,我们试图分析ERT嵌入在马拉雅拉姆语等低资源语言中的仇恨语音检测中的应用。这一领域的研究人员面临的关键挑战是,大多数非英语语言在社交媒体上以代码混合的形式出现。在这里,我们使用基于转换器的模型将tweet分类为仇恨或非仇恨内容。因此,这是一种在非英语文本中使用BERT的新颖方法。
Hate speech detection is a very popular research area for past few years. Hate speech is given various definition by various researchers. In this paper we try to analyse the use of BERT embedding in hate speech detection in low resource language like Malayalam. The crucial challenge faced by researchers in this area are that most non-English languages are represented in code-mixed form in Social media. Here we work with transformer-based models to classify tweets as hate or non-hate content. Hence this is a novel approach that uses BERT in non-English text.