Combining Pre-Trained Language Models and Features for Offensive Language Detection

Combining Pre-Trained Language Models and Features for Offensive Language Detection
复制标题

DOI:
10.1109/iiai-aai-winter58034.2022.00012
复制
发表时间:
2022-12
期刊:
2022 13th International Congress on Advanced Applied Informatics Winter (IIAI-AAI-Winter)
影响因子:
--
通讯作者:
Zhenming Li;Kazutaka Shimada
Zhenming Li;Kazutaka Shimada
中科院分区:
其他
文献类型:
--
作者:
Zhenming Li;Kazutaka Shimada

文献摘要

相似文献

如今,人们经常在社交媒体上更容易地向他人表达他们的辱骂和攻击性想法。辱骂和有毒的评论严重伤害了他人。因此,应该通过自然语言处理来正确检测这些辱骂性和有毒的评论。本文主要讨论冒犯性语言的两类特征:词汇层面特征和言语层面特征。我们使用基于词典和标准的词袋特征作为词级。我们引入基于BERT和基于DeepMoji的特征作为句子级别。我们将这四个特征应用于机器学习方法:支持向量机。我们使用四个特征与一个数据集的组合来评估该方法,好奇猫。最好的F1分数由具有所有特征的方法生成。这一结果表明了我们所提出的方法的有效性。此外,实验结果表明,从Twitter数据生成的DeepMoji是优于BERT,这是从书面语言生成的,对社会媒体数据的攻击性语言检测任务。
Nowadays, people often express their abusive and offensive thoughts to others on social media easier. The abusive and toxic comments hurt others seriously. Therefore those abusive and toxic comments should be detected properly through natural language processing. In this paper, we focus on two types of features in offensive language: word-level and sentence-level fea-tures. We use lexicon-based and standard bag-of-words features as the word level. We introduce BERT-based and DeepMoji-based features as the sentence level. We apply the four features to a machine learning approach: support vector machines. We evaluate the method using the combinations of four features with a dataset, Curious Cat. The best F1 score was generated by the method with all features. This result shows the effectiveness of our proposed method. In addition, the experimental result indicates that DeepMoji generated from Twitter data is better than BERT which is generated from written language, for an offensive language detection task about social media data.