HateCheck: Functional Tests for Hate Speech Detection Models

HateCheck: Functional Tests for Hate Speech Detection Models
复制标题

DOI:
10.18653/v1/2021.acl-long.4
复制
发表时间:
2020-12
期刊:
--
影响因子:
--
通讯作者:
Paul Röttger;B. Vidgen;Dong Nguyen;Zeerak Talat;H. Margetts;J. Pierrehumbert
Paul Röttger;B. Vidgen;Dong Nguyen;Zeerak Talat;H. Margetts;J. Pierrehumbert
中科院分区:
其他
文献类型:
--
作者:
Paul Röttger;B. Vidgen;Dong Nguyen;Zeerak Talat;H. Margetts;J. Pierrehumbert

文献摘要

被引文献

相似文献

检测在线仇恨是一项艰巨的任务,即使是最先进的模型,仇恨言语检测模型是通过使用准确性和F1分数等指标来评估其在Hold-Out测试数据上的艰难方法很难识别特定的模型弱点。由于仇恨言论数据集中越来越好的系统差距和偏见,它也可能高估了可普遍的模型性能。为了启用更多针对性的诊断见解,我们引入了Hatecheck,这是仇恨言论检测模型的功能性测试。我们指定了29个模型功能。通过结构化注释过程来验证其质量,以说明Hatecheck的实用程序作为两个流行的商业模型,揭示了关键模型弱点。
Detecting online hate is a difficult task that even state-of-the-art models struggle with. Typically, hate speech detection models are evaluated by measuring their performance on held-out test data using metrics such as accuracy and F1 score. However, this approach makes it difficult to identify specific model weak points. It also risks overestimating generalisable model performance due to increasingly well-evidenced systematic gaps and biases in hate speech datasets. To enable more targeted diagnostic insights, we introduce HateCheck, a suite of functional tests for hate speech detection models. We specify 29 model functionalities motivated by a review of previous research and a series of interviews with civil society stakeholders. We craft test cases for each functionality and validate their quality through a structured annotation process. To illustrate HateCheck’s utility, we test near-state-of-the-art transformer models as well as two popular commercial models, revealing critical model weaknesses.