Cross-Modal 3D Shape Generation and Manipulation

Cross-Modal 3D Shape Generation and Manipulation
复制标题

DOI:
10.48550/arxiv.2207.11795
复制
发表时间:
2022-07
期刊:
--
影响因子:
--
通讯作者:
Zezhou Cheng;Menglei Chai;Jian Ren;Hsin-Ying Lee;Kyle Olszewski;Zeng Huang;Subhransu Maji;S. Tulyakov
Zezhou Cheng;Menglei Chai;Jian Ren;Hsin-Ying Lee;Kyle Olszewski;Zeng Huang;Subhransu Maji;S. Tulyakov
中科院分区:
其他
文献类型:
--
作者:
Zezhou Cheng;Menglei Chai;Jian Ren;Hsin-Ying Lee;Kyle Olszewski;Zeng Huang;Subhransu Maji;S. Tulyakov

文献摘要

相似文献

创建和编辑 3D 对象的形状和颜色需要大量的人力和专业知识。与 3D 界面中的直接操作相比,草图和涂鸦等 2D 交互对于用户来说通常更加自然和直观。在本文中,我们提出了一种通用的多模态生成模型,该模型通过共享潜在空间将 2D 模态和隐式 3D 表示耦合起来。使用所提出的模型,只需通过潜在空间传播特定 2D 控制模态的编辑即可实现多功能 3D 生成和操作。例如,通过绘制草图来编辑 3D 形状,通过在 2D 渲染上绘制颜色涂鸦来重新着色 3D 表面,或者在给定一个或几个参考图像的情况下生成特定类别的 3D 形状。与之前的工作不同,我们的模型不需要对每个编辑任务进行重新训练或微调,并且概念上简单、易于实现、对输入域转换具有鲁棒性,并且能够灵活地对部分 2D 输入进行不同的重建。我们在灰度线草图和渲染彩色图像的两种代表性 2D 模式上评估我们的框架,并证明我们的方法可以使用这些 2D 模式实现各种形状操作和生成任务。
Creating and editing the shape and color of 3D objects require tremendous human effort and expertise. Compared to direct manipulation in 3D interfaces, 2D interactions such as sketches and scribbles are usually much more natural and intuitive for the users. In this paper, we propose a generic multi-modal generative model that couples the 2D modalities and implicit 3D representations through shared latent spaces. With the proposed model, versatile 3D generation and manipulation are enabled by simply propagating the editing from a specific 2D controlling modality through the latent spaces. For example, editing the 3D shape by drawing a sketch, re-colorizing the 3D surface via painting color scribbles on the 2D rendering, or generating 3D shapes of a certain category given one or a few reference images. Unlike prior works, our model does not require re-training or fine-tuning per editing task and is also conceptually simple, easy to implement, robust to input domain shifts, and flexible to diverse reconstruction on partial 2D inputs. We evaluate our framework on two representative 2D modalities of grayscale line sketches and rendered color images, and demonstrate that our method enables various shape manipulation and generation tasks with these 2D modalities.