Developing an online hate classifier for multiple social media platforms

Developing an online hate classifier for multiple social media platforms
复制标题

DOI:
10.1186/s13673-019-0205-6
复制
发表时间:
2020-01-02
影响因子:
6.6
通讯作者:
Jansen, Bernard J.
Jansen, Bernard J.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Salminen, Joni;Hopf, Maximilian;Jansen, Bernard J.

文献摘要

被引文献

相似文献

社交媒体的普及使人们能够在网上广泛表达自己的意见。然而,与此同时,这也导致了冲突和仇恨的出现,使得在线环境对用户没有吸引力。尽管研究人员发现仇恨是一个跨多个平台的问题,但缺乏使用多平台数据进行在线仇恨检测的模型。为了弥补这一研究空白,我们从YouTube、Reddit、维基百科和Twitter四个平台上收集了197,566条评论,其中80%的评论被标记为非仇恨,其余20%被标记为仇恨。然后,我们实验了几种分类算法(逻辑回归、朴素贝叶斯、支持向量机、XGBoost和神经网络)和特征表示(Bag-of-Words、TF-IDF、Word 2 Vec、BERT及其组合)。虽然所有模型的性能都明显优于基于关键字的基线分类器,但使用所有特征的XGBoost表现最好(F1 = 0.92)。特征重要性分析表明,BERT特征对预测的影响最大。研究结果支持最佳模型的普遍性,因为Twitter和维基百科的特定平台结果与其各自的源文件相当。我们将我们的代码公开用于真实的软件系统中的应用,以及在线仇恨研究人员的进一步开发。
The proliferation of social media enables people to express their opinions widely online. However, at the same time, this has resulted in the emergence of conflict and hate, making online environments uninviting for users. Although researchers have found that hate is a problem across multiple platforms, there is a lack of models for online hate detection using multi-platform data. To address this research gap, we collect a total of 197,566 comments from four platforms: YouTube, Reddit, Wikipedia, and Twitter, with 80% of the comments labeled as non-hateful and the remaining 20% labeled as hateful. We then experiment with several classification algorithms (Logistic Regression, Naive Bayes, Support Vector Machines, XGBoost, and Neural Networks) and feature representations (Bag-of-Words, TF-IDF, Word2Vec, BERT, and their combination). While all the models significantly outperform the keyword-based baseline classifier, XGBoost using all features performs the best (F1 = 0.92). Feature importance analysis indicates that BERT features are the most impactful for the predictions. Findings support the generalizability of the best model, as the platform-specific results from Twitter and Wikipedia are comparable to their respective source papers. We make our code publicly available for application in real software systems as well as for further development by online hate researchers.