Demographic Word Embeddings for Racism Detection on Twitter

Demographic Word Embeddings for Racism Detection on Twitter
复制标题

用于 Twitter 上种族主义检测的人口统计词嵌入

DOI:
--
复制
发表时间:
2017
期刊:
International Joint Conference on Natural Language Processing
影响因子:
--
通讯作者:
Andy Way
Andy Way
中科院分区:
--
文献类型:
--
作者:
Mohammed Hasanuzzaman;G. Dias;Andy Way

文献摘要

被引文献

相似文献

大多数社交媒体平台允许用户自由表达他们的想法、信仰和观点,从而赋予他们言论自由。虽然这代表着难以置信和独特的交流机会,但也带来了重大挑战。网络种族主义就是这样一个例子。在这项研究中,我们提出了一种有监督的学习策略来检测Twitter上的种族主义语言,该策略基于包含人口统计信息(年龄、性别和位置)的单词嵌入。我们的方法在黄金标准数据集(F1=76.3%)上实现了合理的分类精度,并显著改善了人口统计不可知性模型的分类性能。
Most social media platforms grant users freedom of speech by allowing them to freely express their thoughts, beliefs, and opinions. Although this represents incredible and unique communication opportunities, it also presents important challenges. Online racism is such an example. In this study, we present a supervised learning strategy to detect racist language on Twitter based on word embedding that incorporate demographic (Age, Gender, and Location) information. Our methodology achieves reasonable classification accuracy over a gold standard dataset (F1=76.3%) and significantly improves over the classification performance of demographic-agnostic models.