From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models

From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models
复制标题

DOI:
10.48550/arxiv.2305.08283
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Shangbin Feng;Chan Young Park;Yuhan Liu;Yulia Tsvetkov
Shangbin Feng;Chan Young Park;Yuhan Liu;Yulia Tsvetkov
中科院分区:
其他
文献类型:
--
作者:
Shangbin Feng;Chan Young Park;Yuhan Liu;Yulia Tsvetkov

文献摘要

被引文献

相似文献

语言模型(LM)在不同的数据源上进行预训练-新闻,论坛,书籍,在线百科全书。这些数据的很大一部分包括事实和观点,一方面,这些事实和观点庆祝民主和思想的多样性,另一方面,这些事实和观点本身就带有社会偏见。我们的工作开发了新的方法来(1)测量在这样的语料库上训练的LM中的媒体偏见,沿着社会和经济轴,(2)测量在政治偏见的LM上训练的下游NLP模型的公平性。我们专注于仇恨言论和错误信息检测,旨在实证量化预训练数据中的政治(社会,经济)偏见对高风险社会导向任务公平性的影响。我们的研究结果表明,预训练的LM确实有政治倾向,这加强了预训练语料库中存在的两极分化,将社会偏见传播到仇恨言论预测和媒体偏见传播到错误信息检测器中。我们讨论了我们的研究结果对NLP研究的影响,并提出了未来的方向,以减轻不公平。
Language models (LMs) are pretrained on diverse data sources—news, discussion forums, books, online encyclopedias. A significant portion of this data includes facts and opinions which, on one hand, celebrate democracy and diversity of ideas, and on the other hand are inherently socially biased. Our work develops new methods to (1) measure media biases in LMs trained on such corpora, along social and economic axes, and (2) measure the fairness of downstream NLP models trained on top of politically biased LMs. We focus on hate speech and misinformation detection, aiming to empirically quantify the effects of political (social, economic) biases in pretraining data on the fairness of high-stakes social-oriented tasks. Our findings reveal that pretrained LMs do have political leanings which reinforce the polarization present in pretraining corpora, propagating social biases into hate speech predictions and media biases into misinformation detectors. We discuss the implications of our findings for NLP research and propose future directions to mitigate unfairness.