Language-driven synthesis of 3D scenes from scene databases

Language-driven synthesis of 3D scenes from scene databases
复制标题

DOI:
10.1145/3272127.3275035
复制
发表时间:
2018-12
期刊:
ACM Transactions on Graphics (TOG)
影响因子:
--
通讯作者:
Rui Ma;A. Patil;Matthew Fisher;Manyi Li;S. Pirk;Binh-Son Hua;Sai-Kit Yeung;Xin Tong;L. Guibas-L.
Rui Ma;A. Patil;Matthew Fisher;Manyi Li;S. Pirk;Binh-Son Hua;Sai-Kit Yeung;Xin Tong;L. Guibas-L.
中科院分区:
其他
文献类型:
--
作者:
Rui Ma;A. Patil;Matthew Fisher;Manyi Li;S. Pirk;Binh-Son Hua;Sai-Kit Yeung;Xin Tong;L. Guibas-L.

文献摘要

被引文献

相似文献

我们介绍了一种新的框架,使用自然语言来生成和编辑3D室内场景,利用场景语义和文本场景接地从大型注释的3D场景数据库的知识。当在子场景级别执行语义操作时,自然语言编辑界面的优势最强,作用于对象组。我们学习如何通过分析现有的3D场景来操纵这些子场景。我们首先解析来自用户的自然语言命令,并将其转换为语义场景图,该语义场景图用于从与命令匹配的数据库中检索相应的子场景,从而执行编辑。然后,我们通过结合场景上下文可能暗示的其他对象来增强检索到的子场景。最后,通过将增强的子场景与用户的当前场景对齐来合成新的3D场景,其中新对象被拼接到环境中,可能触发对现有场景布置的适当调整。一个具有多种解释的用户命令的暗示性建模界面用于减轻自然语言中的歧义。我们进行研究,比较我们的方法对先前的文本到场景的工作和艺术家制作的场景,发现我们的方法显着优于先前的工作,即使在复杂和不同的自然句子使用的手工制作的场景。
We introduce a novel framework for using natural language to generate and edit 3D indoor scenes, harnessing scene semantics and text-scene grounding knowledge learned from large annotated 3D scene databases. The advantage of natural language editing interfaces is strongest when performing semantic operations at the sub-scene level, acting on groups of objects. We learn how to manipulate these sub-scenes by analyzing existing 3D scenes. We perform edits by first parsing a natural language command from the user and transforming it into a semantic scene graph that is used to retrieve corresponding sub-scenes from the databases that match the command. We then augment this retrieved sub-scene by incorporating other objects that may be implied by the scene context. Finally, a new 3D scene is synthesized by aligning the augmented sub-scene with the user's current scene, where new objects are spliced into the environment, possibly triggering appropriate adjustments to the existing scene arrangement. A suggestive modeling interface with multiple interpretations of user commands is used to alleviate ambiguities in natural language. We conduct studies comparing our approach against both prior text-to-scene work and artist-made scenes and find that our method significantly outperforms prior work and is comparable to handmade scenes even when complex and varied natural sentences are used.