Cyber Hate Speech on Twitter: An Application of Machine Classification and Statistical Modeling for Policy and Decision Making

Cyber Hate Speech on Twitter: An Application of Machine Classification and Statistical Modeling for Policy and Decision Making
复制标题

DOI:
10.1002/poi3.85
复制
发表时间:
2015-06-01
影响因子:
4.9
通讯作者:
Williams, Matthew L.
Williams, Matthew L.
中科院分区:
人文科学2区
文献类型:
--
作者:
Burnap, Pete;Williams, Matthew L.

文献摘要

被引文献

相似文献

在政策和决策中使用“大数据”是当前辩论的主题。2013年鼓手李·里格比在英国伦敦伍尔维奇被谋杀一案在社交媒体上引发了广泛的公众反应,这为研究推特上网络仇恨言论(网络仇恨)的传播提供了机会。在里格比被谋杀后立即收集了人类注释的Twitter数据,以训练和测试监督机器学习文本分类器,该分类器区分仇恨和/或敌对反应,重点是种族,民族或宗教;以及更一般的反应。分类特征来自每条推文的内容,包括单词之间的语法依赖关系,以识别“他者”短语,煽动以对抗行动回应,以及对社会群体有充分理由或正当理由的歧视。分类器的结果是最佳的组合使用概率,基于规则的,基于空间的分类器与表决集成元分类器。我们演示了如何在用于预测Twitter数据样本中网络仇恨可能传播的统计模型中稳健地利用分类器的结果。政策和决策的应用进行了讨论。
The use of "Big Data" in policy and decision making is a current topic of debate. The 2013 murder of Drummer Lee Rigby in Woolwich, London, UK led to an extensive public reaction on social media, providing the opportunity to study the spread of online hate speech (cyber hate) on Twitter. Human annotated Twitter data was collected in the immediate aftermath of Rigby's murder to train and test a supervised machine learning text classifier that distinguishes between hateful and/or antagonistic responses with a focus on race, ethnicity, or religion; and more general responses. Classification features were derived from the content of each tweet, including grammatical dependencies between words to recognize "othering" phrases, incitement to respond with antagonistic action, and claims of well-founded or justified discrimination against social groups. The results of the classifier were optimal using a combination of probabilistic, rule-based, and spatial-based classifiers with a voted ensemble meta-classifier. We demonstrate how the results of the classifier can be robustly utilized in a statistical model used to forecast the likely spread of cyber hate in a sample of Twitter data. The applications to policy and decision making are discussed.