LinkNet: Relational Embedding for Scene Graph

LinkNet: Relational Embedding for Scene Graph
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Sanghyun Woo;Dahun Kim;Donghyeon Cho-;In-So Kweon
Sanghyun Woo;Dahun Kim;Donghyeon Cho-;In-So Kweon
中科院分区:
其他
文献类型:
--
作者:
Sanghyun Woo;Dahun Kim;Donghyeon Cho-;In-So Kweon

文献摘要

被引文献

相似文献

物体及其关系是图像理解的重要内容。场景图提供了捕获图像的这些属性的结构化描述。然而,关于物体之间的关系的推理是非常具有挑战性的,最近只有几个工作试图解决从图像生成场景图的问题。在本文中,我们提出了一种通过显式建模整个对象实例之间的相互依赖来改进场景图生成的方法。我们设计了一个简单有效的关系嵌入模块,使我们的模型能够联合表示所有相关对象之间的联系,而不是孤立地关注一个对象。我们的方法对场景图生成任务的主要部分:关系分类有很大的好处。使用它在基本的更快的R-CNN上,我们的模型在视觉基因组基准上实现了最先进的结果。通过引入全局上下文编码模块和几何布局编码模块,进一步提高了算法的性能。我们通过广泛的烧蚀研究验证了我们最终的模型LinkNet,展示了它在场景图生成中的有效性。
Objects and their relationships are critical contents for image understanding. A scene graph provides a structured description that captures these properties of an image. However, reasoning about the relationships between objects is very challenging and only a few recent works have attempted to solve the problem of generating a scene graph from an image. In this paper, we present a method that improves scene graph generation by explicitly modeling inter-dependency among the entire object instances. We design a simple and effective relational embedding module that enables our model to jointly represent connections among all related objects, rather than focus on an object in isolation. Our method significantly benefits the main part of the scene graph generation task: relationship classification. Using it on top of a basic Faster R-CNN, our model achieves state-of-the-art results on the Visual Genome benchmark. We further push the performance by introducing global context encoding module and geometrical layout encoding module. We validate our final model, LinkNet, through extensive ablation studies, demonstrating its efficacy in scene graph generation.