Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task

Predicting Visual Political Bias Using Webly Supervised Data and an Auxiliary Task
复制标题

DOI:
10.1007/s11263-021-01506-3
复制
发表时间:
2021-08
影响因子:
19.5
通讯作者:
Christopher Thomas;Adriana Kovashka
Christopher Thomas;Adriana Kovashka
中科院分区:
计算机科学2区
文献类型:
--
作者:
Christopher Thomas;Adriana Kovashka

文献摘要

相似文献

新闻媒体塑造公众舆论,对于细心的人类观察者来说,它们所包含的视觉偏见往往是显而易见的。这种偏见可以从不同的媒体来源如何描绘不同的主题或话题中推断出来。在本文中,我们使用网络监督数据大规模模拟当代媒体来源中的视觉政治偏见。我们从左倾和右倾新闻来源收集了超过100万张独特图像和相关新闻文章的数据集,并开发了一种预测图像政治倾向的方法。这个问题特别具有挑战性,因为我们的数据具有巨大的类内视觉和语义多样性。我们建议通过两个阶段的培训来解决这个问题。在第一阶段,模型被迫学习相关的视觉概念,当与从与图像配对的文章中计算的文档嵌入相结合时,使模型能够预测偏差。在第二阶段,我们去除文本域的要求,并从前一个模型的特征中训练一个视觉分类器。我们展示了这种两阶段的方法,它依赖于利用文本的辅助任务,促进学习,并且优于几个强基线。我们提出了广泛的定量和定性结果分析我们的数据集。我们的研究结果揭示了政治光谱的不同方面如何描绘个人、群体和主题的差异。
The news media shape public opinion, and often, the visual bias they contain is evident for careful human observers. This bias can be inferred from how different media sources portray different subjects or topics. In this paper, we model visual political bias in contemporary media sources at scale, using webly supervised data. We collect a dataset of over one million unique images and associated news articles from left- and right-leaning news sources, and develop a method to predict the image’s political leaning. This problem is particularly challenging because of the enormous intra-class visual and semantic diversity of our data. We propose two stages of training to tackle this problem. In the first stage, the model is forced to learn relevant visual concepts that, when joined with document embeddings computed from articles paired with the images, enable the model to predict bias. In the second stage, we remove the requirement of the text domain and train a visual classifier from the features of the former model. We show this two-stage approach that relies on an auxiliary task leveraging text, facilitates learning and outperforms several strong baselines. We present extensive quantitative and qualitative results analyzing our dataset. Our results reveal disparities in how different sides of the political spectrum portray individuals, groups, and topics.