Tackling racial bias in automated online hate detection: Towards fair and accurate detection of hateful users with geometric deep learning

Tackling racial bias in automated online hate detection: Towards fair and accurate detection of hateful users with geometric deep learning
复制标题

DOI:
10.1140/epjds/s13688-022-00319-9
复制
发表时间:
2022-02
期刊:
影响因子:
3.6
通讯作者:
Zo Ahmed;Bertie Vidgen;Scott A. Hale
Zo Ahmed;Bertie Vidgen;Scott A. Hale
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zo Ahmed;Bertie Vidgen;Scott A. Hale

文献摘要

被引文献

相似文献

在许多社交媒体平台上,网络仇恨日益受到关注,这让他们变得不受欢迎和不安全。为了应对这一问题,科技公司正在越来越多地开发技术,以自动识别和制裁讨厌的用户。然而,由于语音的语境性质,对这类用户的准确检测仍然是一个挑战,其含义取决于使用它的社会环境。言论的这种语境性质也导致了小型化的用户,特别是非裔美国人,被设计来保护他们的算法不公平地检测为“可恨的”。为了解决这个不准确和不公平的仇恨检测问题,研究集中在开发更好地理解文本上下文的机器学习(ML)系统。尽管社会科学研究表明,整合可恶用户的社交网络提供了丰富的背景信息,但它并没有得到如此多的关注。提出了一种通过几何深度学习结合社交网络信息来更准确、更公平地检测恶意用户的系统。几何深度学习是一种动态学习信息丰富的网络表示的ML技术。我们的主要贡献有两个:首先,我们证明了与其他完全排除网络信息或通过人工特征工程合并网络信息的方法相比,使用几何深度学习来添加网络信息可以产生更准确的分类器。我们的最佳性能模型在之前发布的可恨用户数据集上获得了90.8%的AUC分数。其次,我们证明了这样的信息也会导致更公平的结果:使用‘预测相等’的公平标准,我们将我们的几何学习算法的假阳性率与其他ML技术进行比较,发现我们性能最好的分类器在非裔美国人用户子集中没有假阳性。没有网络信息的神经网络的假阳性数最多,为26个,而包含人工网络功能的神经网络在非裔美国人用户中有13个假阳性。我们介绍的系统强调了有效地将社交网络功能纳入自动仇恨用户检测的重要性,为改善如何解决在线仇恨提供了新的机会。
Online hate is a growing concern on many social media platforms, making them unwelcoming and unsafe. To combat this, technology companies are increasingly developing techniques to automatically identify and sanction hateful users. However, accurate detection of such users remains a challenge due to the contextual nature of speech, whose meaning depends on the social setting in which it is used. This contextual nature of speech has also led to minoritized users, especially African–Americans, to be unfairly detected as ‘hateful’ by the very algorithms designed to protect them. To resolve this problem of inaccurate and unfair hate detection, research has focused on developing machine learning (ML) systems that better understand textual context. Incorporating social networks of hateful users has not received as much attention, despite social science research suggesting it provides rich contextual information. We present a system for more accurately and fairly detecting hateful users by incorporating social network information through geometric deep learning. Geometric deep learning is a ML technique that dynamically learns information-rich network representations. We make two main contributions: first, we demonstrate that adding network information with geometric deep learning produces a more accurate classifier compared with other techniques that either exclude network information entirely or incorporate it through manual feature engineering. Our best performing model achieves an AUC score of 90.8% on a previously released hateful user dataset. Second, we show that such information also leads to fairer outcomes: using the ‘predictive equality’ fairness criteria, we compare the false positive rates of our geometric learning algorithm to other ML techniques and find that our best-performing classifier has no false positives among a subset of African–American users. A neural network without network information has the largest number of false positives at 26, while a neural network incorporating manual network features has 13 false positives among African–American users. The system we present highlights the importance of effectively incorporating social network features in automated hateful user detection, raising new opportunities to improve how online hate is tackled.