A literature survey on multimodal and multilingual automatic hate speech identification

A literature survey on multimodal and multilingual automatic hate speech identification
复制标题

DOI:
10.1007/s00530-023-01051-8
复制
发表时间:
2023-01
期刊:
影响因子:
3.9
通讯作者:
Anusha Chhabra;D. Vishwakarma
Anusha Chhabra;D. Vishwakarma
中科院分区:
计算机科学4区
文献类型:
--
作者:
Anusha Chhabra;D. Vishwakarma

文献摘要

被引文献

相似文献

社交媒体是一个更常见和强大的交流平台,可以分享关于任何主题或文章的观点,从而导致非结构化的有毒和仇恨的对话。遏制仇恨言论已成为全球面临的一项重大挑战。在这方面,社交媒体平台正在使用人工智能技术的现代统计工具来处理和消除有毒数据,以最大限度地减少全球仇恨犯罪。由于迫切的需求,基于机器和深度学习的技术在分析这类数据时越来越受到关注。这项调查全面分析了仇恨言论的定义,沿着分析了检测的动机以及在识别仇恨言论方面发挥关键作用的标准文本分析方法。还讨论了最先进的仇恨言论识别方法,通过考虑多模式和多语言输入,强调了手工制作的基于特征和基于深度学习的算法,并说明了每种方法的优缺点。调查还提出了流行的基准数据集仇恨言论/攻击性语言检测指定他们的挑战,实现最高分类分数的方法,以及数据集的特征,如样本数量,模态,语言(S),类的数量等,此外,性能指标进行了描述,并提到流行的仇恨言论方法的分类分数。最后给出了研究结论和未来的研究方向。与早期的调查相比,本文通过组织良好的比较,挑战和最新的评估技术,沿着他们的最佳表现,更好地介绍了多模式和多语言仇恨言论检测。
Social media is a more common and powerful platform for communication to share views about any topic or article, which consequently leads to unstructured toxic, and hateful conversations. Curbing hate speeches has emerged as a critical challenge globally. In this regard, Social media platforms are using modern statistical tools of AI technologies to process and eliminate toxic data to minimize hate crimes globally. Demanding the dire need, machine and deep learning-based techniques are getting more attention in analyzing these kinds of data. This survey presents a comprehensive analysis of hate speech definitions along with the motivation for detection and standard textual analysis methods that play a crucial role in identifying hate speech. State-of-the-art hate speech identification methods are also discussed, highlighting handcrafted feature-based and deep learning-based algorithms by considering multimodal and multilingual inputs and stating the pros and cons of each. Survey also presents popular benchmark datasets of hate speech/offensive language detection specifying their challenges, the methods for achieving top classification scores, and dataset characteristics such as the number of samples, modalities, language(s), number of classes, etc. Additionally, performance metrics are described, and classification scores of popular hate speech methods are mentioned. The conclusion and future research directions are presented at the end of the survey. Compared with earlier surveys, this paper gives a better presentation of multimodal and multilingual hate speech detection through well-organized comparisons, challenges, and the latest evaluation techniques, along with their best performances.