Hypernymization of named entity-rich captions for grounding-based multi-modal pretraining

Hypernymization of named entity-rich captions for grounding-based multi-modal pretraining
复制标题

DOI:
10.1145/3591106.3592223
复制
发表时间:
2023-04
期刊:
Proceedings of the 2023 ACM International Conference on Multimedia Retrieval
影响因子:
--
通讯作者:
Giacomo Nebbia;Adriana Kovashka
Giacomo Nebbia;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Giacomo Nebbia;Adriana Kovashka

文献摘要

相似文献

命名实体在自然伴随图像的文本中无处不在,特别是在新闻或维基百科文章等领域。在以前的工作中,命名实体已被确定为在维基百科上预训练并在无命名实体基准数据集上评估的图像-文本检索模型性能低下的可能原因。由于命名实体很少被提及,因此它们的建模可能具有挑战性。它们也代表了自监督模型错过的学习机会:模型可能会错过图像中命名实体和对象之间的联系,但如果使用更常见的术语提到对象,则不会。在这项工作中,我们研究hypernymization作为一种处理命名实体的方法,用于预训练基于接地的多模态模型和对开放词汇检测进行微调。我们提出了两种方法来执行hypernymization:(1)依赖于概念的综合本体的“手动”管道,以及(2)“学习”方法,其中我们训练语言模型来学习执行hypernymization。我们对维基百科和《纽约时报》的数据进行了实验。我们报告了hypernymization后感兴趣对象的预训练性能的改善,并且我们展示了hypernymization对开放词汇检测的承诺,特别是在训练过程中看不到的类上。
Named entities are ubiquitous in text that naturally accompanies images, especially in domains such as news or Wikipedia articles. In previous work, named entities have been identified as a likely reason for low performance of image-text retrieval models pretrained on Wikipedia and evaluated on named entities-free benchmark datasets. Because they are rarely mentioned, named entities could be challenging to model. They also represent missed learning opportunities for self-supervised models: the link between named entity and object in the image may be missed by the model, but it would not be if the object were mentioned using a more common term. In this work, we investigate hypernymization as a way to deal with named entities for pretraining grounding-based multi-modal models and for fine-tuning on open-vocabulary detection. We propose two ways to perform hypernymization: (1) a “manual” pipeline relying on a comprehensive ontology of concepts, and (2) a “learned” approach where we train a language model to learn to perform hypernymization. We run experiments on data from Wikipedia and from The New York Times. We report improved pretraining performance on objects of interest following hypernymization, and we show the promise of hypernymization on open-vocabulary detection, specifically on classes not seen during training.