Unsupervised Visual Relationship Inference

Unsupervised Visual Relationship Inference
复制标题

无监督视觉关系推理

DOI:
10.1109/icip40778.2020.9190770
复制
发表时间:
2020
期刊:
Proceedings of the IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Hideki Nakayama
Hideki Nakayama
中科院分区:
--
文献类型:
--
作者:
Taiga Kashima;Kento Masui;Hideki Nakayama

文献摘要

相似文献

视觉关系推理是图像理解的一个重要研究领域。由于最近深度学习的进展,在这一具有挑战性的领域取得了重大进展的迹象。标准方法试图通过使用仔细注释的数据集来基于监督学习来识别视觉关系,其中附加了图像、三元组(主语-谓语-宾语)和边界框。然而,准备大规模数据集是非常耗时的。这项研究提出了一种新的方法来推断视觉关系,而不需要图像-三元组对。我们的方法试图保持所推断的三元组的循环一致性和似然性。我们的实验结果表明,该方法能够在未配对的环境中推断对象之间的谓词,并且使用从外部图像描述中解析的三元组也取得了令人满意的结果。
Visual relationship inference is an essential research area for image understanding. Owing to the recent advancement of deep learning, significant signs of progress have been made in this challenging area. Standard approaches attempt to recognize visual relationships based on supervised learning by employing a carefully annotated dataset, in which images, triplets (subject-predicate-object), and bounding boxes are attached. However, preparing a large-scale dataset is very time consuming. This study proposes a novel method to infer visual relationships without image-triplet pairs. Our method tries to keep cycle consistency and plausibility of the inferred triplets. Our experimental results demonstrate that this method can infer predicates between objects in unpaired settings, and also achieving promising results using triplets parsed from external image descriptions.