Unconditional Scene Graph Generation

Unconditional Scene Graph Generation
复制标题

DOI:
10.1109/iccv48922.2021.01605
复制
发表时间:
2021-08
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Sarthak Garg;Helisa Dhamo;Azade Farshad;S. Musatian;Nassir Navab;F. Tombari
Sarthak Garg;Helisa Dhamo;Azade Farshad;S. Musatian;Nassir Navab;F. Tombari
中科院分区:
其他
文献类型:
--
作者:
Sarthak Garg;Helisa Dhamo;Azade Farshad;S. Musatian;Nassir Navab;F. Tombari

文献摘要

相似文献

尽管最近在单域或单对象图像生成方面取得了进展,但生成包含不同、多个对象及其交互的复杂场景仍然具有挑战性。场景图由作为对象的节点和作为对象之间关系的有向边组成,提供了比图像更具有语义基础的场景的替代表示。我们假设场景图的生成模型可能能够比图像更有效地学习现实世界场景的底层语义结构,从而以场景图的形式生成真实的新颖场景。在这项工作中,我们探索了无条件生成语义场景图的新任务。我们开发了一种名为 SceneGraphGen 的深度自回归模型,它可以使用分层循环架构直接学习标记图和有向图上的概率分布。该模型以种子对象作为输入,并按一系列步骤生成场景图,每个步骤生成一个对象节点,后面是连接到先前节点的一系列关系边。我们证明了 SceneGraphGen 生成的场景图是多种多样的,并且遵循现实世界场景的语义模式。此外,我们还演示了生成的图在图像合成、异常检测和场景图补全中的应用。
Despite recent advancements in single-domain or single-object image generation, it is still challenging to generate complex scenes containing diverse, multiple objects and their interactions. Scene graphs, composed of nodes as objects and directed-edges as relationships among objects, offer an alternative representation of a scene that is more semantically grounded than images. We hypothesize that a generative model for scene graphs might be able to learn the underlying semantic structure of real-world scenes more effectively than images, and hence, generate realistic novel scenes in the form of scene graphs. In this work, we explore a new task for the unconditional generation of semantic scene graphs. We develop a deep auto-regressive model called SceneGraphGen which can directly learn the probability distribution over labelled and directed graphs using a hierarchical recurrent architecture. The model takes a seed object as input and generates a scene graph in a sequence of steps, each step generating an object node, followed by a sequence of relationship edges connecting to the previous nodes. We show that the scene graphs generated by SceneGraphGen are diverse and follow the semantic patterns of real-world scenes. Additionally, we demonstrate the application of the generated graphs in image synthesis, anomaly detection and scene graph completion.