Linguistic Structures as Weak Supervision for Visual Scene Graph Generation

Linguistic Structures as Weak Supervision for Visual Scene Graph Generation
复制标题

DOI:
10.1109/cvpr46437.2021.00819
复制
发表时间:
2021-05
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Keren Ye;Adriana Kovashka
Keren Ye;Adriana Kovashka
中科院分区:
其他
文献类型:
--
作者:
Keren Ye;Adriana Kovashka

文献摘要

相似文献

之前在场景图生成方面的工作需要在三联体水平上进行分类监督——主体和客体,以及与它们相关的谓词,有或没有边界框信息。然而,场景图生成是一个整体任务:因此,整体的上下文监督应该直观地提高性能。在这项工作中,我们探讨了字幕中的语言结构如何有利于场景图的生成。我们的方法捕获标题中提供的关于单个三元组之间关系的信息,以及主题和对象的上下文(例如提到的视觉属性)。标题是一种弱于三联体的监督类型,因为三联体中人类注释的主题和对象的详尽列表与标题中的名词之间的对齐是弱的。然而,考虑到网络上多模态数据的大量和多样化来源(例如带有图像和说明文字的博客文章),语言监督比众包三元组更具可扩展性。我们展示了与先前利用实例级和图像级监督的方法进行了广泛的实验比较,并简化了我们的方法,以显示利用短语和顺序上下文的影响,以及改进主体和客体定位的技术。
Prior work in scene graph generation requires categorical supervision at the level of triplets—subjects and objects, and predicates that relate them, either with or without bounding box information. However, scene graph generation is a holistic task: thus holistic, contextual supervision should intuitively improve performance. In this work, we explore how linguistic structures in captions can benefit scene graph generation. Our method captures the information provided in captions about relations between individual triplets, and context for subjects and objects (e.g. visual properties are mentioned). Captions are a weaker type of supervision than triplets since the alignment between the exhaustive list of human-annotated subjects and objects in triplets, and the nouns in captions, is weak. However, given the large and diverse sources of multimodal data on the web (e.g. blog posts with images and captions), linguistic supervision is more scalable than crowdsourced triplets. We show extensive experimental comparisons against prior methods which leverage instance- and image-level supervision, and ablate our method to show the impact of leveraging phrasal and sequential context, and techniques to improve localization of subjects and objects.