Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis

Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis
复制标题

DOI:
10.48550/arxiv.2211.02269
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Changyuan Qiu;Winston Wu;Xinliang Frederick Zhang;Lu Wang
Changyuan Qiu;Winston Wu;Xinliang Frederick Zhang;Lu Wang
中科院分区:
其他
文献类型:
--
作者:
Changyuan Qiu;Winston Wu;Xinliang Frederick Zhang;Lu Wang

文献摘要

相似文献

先前关于意识形态预测的工作主要集中在单一模式,即文本或图像。在这项工作中,我们介绍了多模态意识形态预测的任务,其中模型在给定具有政治内容的文本图像对的情况下预测二元或五点尺度的意识形态倾向。我们首先收集了五个新的大型数据集,其中包含英文文档和图像及其意识形态倾向,涵盖美国各种主流媒体的新闻文章以及 Reddit 和 Twitter 的社交媒体帖子。我们对新闻文章进行深入分析,揭示不同政治派别在图像内容和使用方面的差异。此外,我们进行了广泛的实验和消融研究,证明了针对不同模型组件的有针对性的预训练目标的有效性。我们性能最好的模型是一种后期融合架构,经过多模态内容的三重目标进行预训练,其性能比最先进的纯文本模型高出近 4%,并且比没有预训练的强大多模态基线高出超过 3%。
Prior work on ideology prediction has largely focused on single modalities, i.e., text or images. In this work, we introduce the task of multimodal ideology prediction, where a model predicts binary or five-point scale ideological leanings, given a text-image pair with political content. We first collect five new large-scale datasets with English documents and images along with their ideological leanings, covering news articles from a wide range of mainstream media in US and social media posts from Reddit and Twitter. We conduct in-depth analyses on news articles and reveal differences in image content and usage across the political spectrum. Furthermore, we perform extensive experiments and ablation studies, demonstrating the effectiveness of targeted pretraining objectives on different model components. Our best-performing model, a late-fusion architecture pretrained with a triplet objective over multimodal content, outperforms the state-of-the-art text-only model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.