Scene Graph Prediction with Limited Labels

Scene Graph Prediction with Limited Labels
复制标题

DOI:
10.1109/iccvw.2019.00220
复制
发表时间:
2019-04
期刊:
2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)
影响因子:
--
通讯作者:
V. Chen;P. Varma;Ranjay Krishna;Michael S. Bernstein;Christopher Ré;Li Fei-Fei-Li-Fei-Fei-48004138
V. Chen;P. Varma;Ranjay Krishna;Michael S. Bernstein;Christopher Ré;Li Fei-Fei-Li-Fei-Fei-48004138
中科院分区:
其他
文献类型:
--
作者:
V. Chen;P. Varma;Ranjay Krishna;Michael S. Bernstein;Christopher Ré;Li Fei-Fei-Li-Fei-Fei-48004138

文献摘要

被引文献

相似文献

视觉知识库(如Visual Genome)为计算机视觉领域的许多应用提供了动力,包括视觉问答和字幕,但存在稀疏、不完整的关系。到目前为止,所有的场景图模型都局限于训练一小组视觉关系,每个视觉关系都有数千个训练标签。雇用人工注释者是昂贵的,并且使用文本知识库补全方法与可视化数据不兼容。在本文中,我们引入了一种半监督方法,该方法使用很少的标记样本为大量未标记的图像分配概率关系标签。我们分析了视觉关系,提出了两种类型的图像不可知特征,用于生成噪声启发式,其输出使用基于因子图的生成模型进行聚合。每个关系只需10个标记示例,生成模型就可以创建足够的训练数据来训练任何现有的最先进的场景图模型。我们证明,我们的方法在PREDCLS的场景图预测上优于所有基线方法5.16 recall@100。在我们的有限标签设置中,我们为关系定义了一个复杂性度量,作为我们的方法优于迁移学习的条件的指标(R^2 = 0.778),迁移学习是使用有限标签进行训练的实际方法。
Visual knowledge bases such as Visual Genome power numerous applications in computer vision, including visual question answering and captioning, but suffer from sparse, incomplete relationships. All scene graph models to date are limited to training on a small set of visual relationships that have thousands of training labels each. Hiring human annotators is expensive, and using textual knowledge base completion methods are incompatible with visual data. In this paper, we introduce a semi-supervised method that assigns probabilistic relationship labels to a large number of unlabeled images using few labeled examples. We analyze visual relationships to suggest two types of image-agnostic features that are used to generate noisy heuristics, whose outputs are aggregated using a factor graph-based generative model. With as few as 10 labeled examples per relationship, the generative model creates enough training data to train any existing state-of-the-art scene graph model. We demonstrate that our method outperforms all baseline approaches on scene graph prediction by5.16 recall@100 for PREDCLS. In our limited label setting, we define a complexity metric for relationships that serves as an indicator (R^2 = 0.778) for conditions under which our method succeeds over transfer learning, the de-facto approach for training with limited labels.