Machine Learning Approach for the Detection of Hate Speech in Sinhala Unicode Text
Machine Learning Approach for the Detection of Hate Speech in Sinhala Unicode Text
复制标题
用于检测僧伽罗 Unicode 文本中仇恨言论的机器学习方法
DOI:
10.1109/icter51097.2020.9325493
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
M. Punchimudiyanse
中科院分区:
文献类型:
--
作者:
S. Samarasinghe;R. Meegama;M. Punchimudiyanse
Hate speech published online platforms has become a critical issue in Sri Lanka since this has caused conflicts between different ethnic groups. One of the main barriers to stop this crime is the lack of resources to detect online hate content in Sinhala automatically. Due to the vast amount of content published on online platforms every minute, an automatic method must be implemented in order to solve this issue.As a solution, we suggest a deep learning mechanism that utilizes two convolution neural networks (CNNs) which will first classify a given text corpus as hateful or not. Then, if the text corpus contains hate content text, it will again be classified according to its hate level which can be used by authorities to make decisions. In order to convert the text data into numerical vectors, we have used FastText word embedding in this study.Results indicate an accuracy of 83% and 60% for hate speech classification and hate level classifications, respectively.