TextDeformer: Geometry Manipulation using Text Guidance

TextDeformer: Geometry Manipulation using Text Guidance
复制标题

DOI:
10.1145/3588432.3591552
复制
发表时间:
2023-04
期刊:
ACM SIGGRAPH 2023 Conference Proceedings
影响因子:
--
通讯作者:
William Gao;Noam Aigerman;Thibault Groueix;Vladimir G. Kim;Rana Hanocka
William Gao;Noam Aigerman;Thibault Groueix;Vladimir G. Kim;Rana Hanocka
中科院分区:
其他
文献类型:
--
作者:
William Gao;Noam Aigerman;Thibault Groueix;Vladimir G. Kim;Rana Hanocka

文献摘要

被引文献

相似文献

我们提出了一种技术,用于自动产生的输入三角形网格的变形,仅由文本提示。我们的框架能够产生大的低频形状变化和小的高频细节的变形。我们的框架依赖于可微分渲染来将几何体连接到强大的预训练图像编码器,如CLIP和DINO。值得注意的是,通过可微分渲染采取梯度步骤来更新网格几何形状是非常具有挑战性的,通常会导致变形的网格具有显著的伪影。这些困难被CLIP的噪声和不一致的梯度放大。为了克服这一限制,我们选择通过雅可比矩阵来表示网格变形,它以全局平滑的方式更新变形(而不是局部次优步骤)。我们的关键观察是,雅可比矩阵是一种有利于更平滑,大变形的表示,导致顶点和像素之间的全局关系,并避免局部噪声梯度。此外,为了确保从所有3D视点得到的形状是一致的,我们鼓励在渲染的2D编码上计算的深度特征对于来自所有视点的给定顶点是一致的。我们证明了我们的方法能够平滑变形各种各样的源网格和目标文本提示,实现大的修改,例如,动物的身体比例,以及添加精细的语义细节,如军队靴子上的鞋带和面部的精细细节。
We present a technique for automatically producing a deformation of an input triangle mesh, guided solely by a text prompt. Our framework is capable of deformations that produce both large, low-frequency shape changes, and small high-frequency details. Our framework relies on differentiable rendering to connect geometry to powerful pre-trained image encoders, such as CLIP and DINO. Notably, updating mesh geometry by taking gradient steps through differentiable rendering is notoriously challenging, commonly resulting in deformed meshes with significant artifacts. These difficulties are amplified by noisy and inconsistent gradients from CLIP. To overcome this limitation, we opt to represent our mesh deformation through Jacobians, which updates deformations in a global, smooth manner (rather than locally-sub-optimal steps). Our key observation is that Jacobians are a representation that favors smoother, large deformations, leading to a global relation between vertices and pixels, and avoiding localized noisy gradients. Additionally, to ensure the resulting shape is coherent from all 3D viewpoints, we encourage the deep features computed on the 2D encoding of the rendering to be consistent for a given vertex from all viewpoints. We demonstrate that our method is capable of smoothly-deforming a wide variety of source mesh and target text prompts, achieving both large modifications to, e.g., body proportions of animals, as well as adding fine semantic details, such as shoe laces on an army boot and fine details of a face.