ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction

ViSER: Video-Specific Surface Embeddings for Articulated 3D Shape Reconstruction
复制标题

ViSER:用于铰接 3D 形状重建的视频特定表面嵌入

DOI:
--
复制
发表时间:
2021
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Deva Ramanan
Deva Ramanan
中科院分区:
--
文献类型:
--
作者:
Gengshan Yang;Deqing Sun;Varun Jampani;Daniel Vlasic;Forrester Cole;Ce Liu;Deva Ramanan

文献摘要

被引文献

相似文献

我们介绍了ViSER,一种从单目视频中恢复铰接3D形状和密集3D轨迹的方法。以前对动态3D形状的高质量重建工作通常依赖于多个摄像机视图、强类别特定先验或2D关键点监督。我们表明,如果可以可靠地估计视频中的远程对应,则不需要这些,仅使用2D对象掩模和两帧光流作为输入。ViSER通过捕捉每个表面点的像素外观的视频特定表面嵌入,将2D像素与规范的可变形3D网格相匹配,从而推断出对应关系。这些嵌入表现为在网格表面上定义的一组连续的关键点描述符,可用于在像素之间建立密集的远程对应。表面嵌入被实现为基于坐标的mlp,通过一致性和对比重建损失来适合每个视频。实验结果表明,与DAVIS和YTVOS的具有挑战性的视频(穿着宽松衣服和不寻常姿势的人)以及动物视频相比,ViSER的效果更好。我们的代码可以在viser-shape.github上获得。io。
We introduce ViSER, a method for recovering articulated 3D shapes and dense 3D trajectories from monocular videos. Previous work on high-quality reconstruction of dynamic 3D shapes typically relies on multiple camera views, strong category-specific priors, or 2D keypoint supervision. We show that none of these are required if one can reliably estimate long-range correspondences in a video, making use of only 2D object masks and two-frame optical flow as inputs. ViSER infers correspondences by matching 2D pixels to a canonical, deformable 3D mesh via video-specific surface embeddings that capture the pixel appearance of each surface point. These embeddings behave as a continuous set of keypoint descriptors defined over the mesh surface, which can be used to establish dense long-range correspondences across pixels. The surface embeddings are implemented as coordinate-based MLPs that are fit to each video via consistency and contrastive reconstruction losses. Experimental results show that ViSER compares favorably against prior work on challenging videos of humans with loose clothing and unusual poses as well as animals videos from DAVIS and YTVOS. Our code is available at viser-shape.github.io .