Machine Learning Approach for the Detection of Hate Speech in Sinhala Unicode Text

Machine Learning Approach for the Detection of Hate Speech in Sinhala Unicode Text
复制标题

用于检测僧伽罗 Unicode 文本中仇恨言论的机器学习方法

DOI:
10.1109/icter51097.2020.9325493
复制
发表时间:
2020
期刊:
2020 20th International Conference on Advances in ICT for Emerging Regions (ICTer)
影响因子:
--
通讯作者:
M. Punchimudiyanse
M. Punchimudiyanse
中科院分区:
--
文献类型:
--
作者:
S. Samarasinghe;R. Meegama;M. Punchimudiyanse

文献摘要

被引文献

相似文献

在斯里兰卡,在网上发表的仇恨言论已成为一个关键问题,因为这引起了不同族裔群体之间的冲突。阻止这种犯罪的主要障碍之一是缺乏资源来自动检测僧伽罗语的在线仇恨内容。由于每分钟在线平台上发布的内容数量巨大,因此必须实现自动方法来解决这个问题。作为解决方案,我们提出了一种利用两个卷积神经网络(CNN)的深度学习机制,它首先将给定的文本语料库分类为仇恨或非仇恨。然后,如果文本语料库包含仇恨内容文本,则它将再次根据其仇恨级别进行分类,这可以由当局用于做出决策。为了将文本数据转换为数字向量,我们在这项研究中使用FastText单词嵌入。结果表明,仇恨言论分类和仇恨水平分类的准确率分别为83%和60%。
Hate speech published online platforms has become a critical issue in Sri Lanka since this has caused conflicts between different ethnic groups. One of the main barriers to stop this crime is the lack of resources to detect online hate content in Sinhala automatically. Due to the vast amount of content published on online platforms every minute, an automatic method must be implemented in order to solve this issue.As a solution, we suggest a deep learning mechanism that utilizes two convolution neural networks (CNNs) which will first classify a given text corpus as hateful or not. Then, if the text corpus contains hate content text, it will again be classified according to its hate level which can be used by authorities to make decisions. In order to convert the text data into numerical vectors, we have used FastText word embedding in this study.Results indicate an accuracy of 83% and 60% for hate speech classification and hate level classifications, respectively.