Reduced-Bias Co-trained Ensembles for Weakly Supervised Cyberbullying Detection

Reduced-Bias Co-trained Ensembles for Weakly Supervised Cyberbullying Detection
复制标题

用于弱监督网络欺凌检测的减少偏差联合训练集合

DOI:
10.1007/978-3-030-34980-6_32
复制
发表时间:
2019
期刊:
2019 IEEE Fifth International Conference on Multimedia Big Data (BigMM)
影响因子:
--
通讯作者:
Bert Huang
Bert Huang
中科院分区:
--
文献类型:
--
作者:
Elaheh Raisi;Bert Huang

文献摘要

被引文献

相似文献

社交媒体反映了社会的许多方面,包括基于性别、种族、宗教、身体能力和性取向等敏感特征对个人的社会偏见。因此,在社交媒体数据上训练的机器学习算法可能会延续或放大对各种人口群体的歧视态度,导致不公平的决策。机器学习的一个重要应用是自动检测网络欺凌。在这种情况下,偏见可能采取欺凌检测器的形式,对某些身份群体或关于某些身份群体的消息进行更频繁的错误检测。在本文中,我们提出了一种从弱监督中训练欺凌检测器的方法,同时降低了学习模型反映或放大数据中歧视性偏见的程度。我们的目标是降低模型对描述特定社会群体的语言的敏感性。一个理想的、公平的基于语言的检测器应该公平地对待描述特定社会群体亚群的语言。建立在以前提出的弱监督学习算法,我们惩罚的模型时,歧视观察。通过惩罚不公平,我们鼓励学习算法在其预测中避免不公平行为,并实现对受保护子群体的公平对待。我们引入了两个不公平的惩罚条款:一个是针对移除公平,另一个是针对替代公平。我们定量和定性地评估所产生的模型的公平性的合成基准和数据从Twitter比较众包注释。
Social media reflects many aspects of society, including social biases against individuals based on sensitive characteristics such as gender, race, religion, physical ability, and sexual orientation. Machine learning algorithms trained on social media data may therefore perpetuate or amplify discriminatory attitudes against various demographic groups, causing unfair decision-making. One important application for machine learning is the automatic detection of cyberbullying. Biases in this context could take the form of bullying detectors that make false detections more frequently on messages by or about certain identity groups. In this paper, we present an approach for training bullying detectors from weak supervision while reducing the degree to which learned models reflect or amplify discriminatory biases in the data. Our goal is to decrease the sensitivity of models to language describing particular social groups. An ideal, fair language-based detector should treat language describing subpopulations of particular social groups equitably. Building on a previously proposed weakly supervised learning algorithm, we penalize the model when discrimination is observed. By penalizing unfairness, we encourage the learning algorithm to avoid unfair behavior in its predictions and achieve equitable treatment for protected subpopulations. We introduce two unfairness penalty terms: one aimed at removal fairness and another at substitutional fairness. We quantitatively and qualitatively evaluate the resulting models’ fairness on a synthetic benchmark and data from Twitter comparing against crowdsourced annotation.