Supplementary of Self-Supervised 3D Mesh Reconstruction from Single Images

Supplementary of Self-Supervised 3D Mesh Reconstruction from Single Images
复制标题

单图像自监督 3D 网格重建的补充

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
Jiaya Jia
Jiaya Jia
中科院分区:
--
文献类型:
--
作者:
T. Hu;Liwei Wang;Xiaogang Xu;Shu Liu;Jiaya Jia

文献摘要

被引文献

相似文献

1.网络结构输入数据X = [I,M ] ∈ RH×W×4由原始RGB单图像I ∈ RH×W×3及其轮廓掩模M ∈ RH×W×1连接而成。我们的3D网格重建模型由4个子编码器组成,即Camera Encoder Ec,Light Encoder El,Shape Encoder Es和Texture Encoder Et。这些子编码器分别对相应的属性特征进行编码,以避免相互影响,如图1所示。地标特征提取器Ef是传统的U-Net网络,其输出分辨率等于输入分辨率H ×W,输出通道为128。地标分类网络是一个输出维数为V的三层MLP网络。2.局限性和失败案例虽然我们的SMR可以在多个类别特定的数据集上有效地重建3D网格对象,只有2D轮廓注释,但在不久的将来仍然有两个主要的局限性需要克服:局限性1:轮廓注释。它仍然需要轮廓注释。为了简化起见,我们没有考虑背景信息的影响。对于一些特定类别的对象,例如马,我们通过detectron 2对其轮廓注释进行注释。因此,不正确的注释将影响重建结果。故障情况如图2所示。限制2:非刚性物体。它不适合非网格对象,如人或花。由于很难确定网格表示的规范视点和拓扑限制。为了避免这些限制,我们将通过自监督学习进一步预测轮廓掩模,并为更一般的对象构建完全无监督的3D重建模型。3. ShapeNet上的3D重建在ShapeNet数据集上,由于我们的SMR旨在对特定类别的对象进行建模,因此我们对这13个类别逐一进行实验。我们引入了地面实况摄像机参数来评估重建精度,并与其他监督方法进行了比较,以证明SMR的有效性5x 5 Conv,32 5x 5 Conv,64 5x 5 Conv,128
1. Network Structure The input data X = [I,M ] ∈ RH×W×4 is concatenated by the original RGB single image I ∈ RH×W×3 and its silhouette mask M ∈ RH×W×1. Our 3D mesh reconstruction model consists of 4 sub-encoders, i.e. Camera Encoder Ec, Light Encoder El, Shape Encoder Es, and Texture Encoder Et. These sub-encoders separately encode corresponding attribute feature to avoid mutual effect, as illustrated in Fig. 1. The landmark feature extractor Ef is the conventional U-Net network whose output resolution is equal to the input resolution H ×W and output channel is 128. The landmark classification network is a MLP network with three layers whose output dimension is V . 2. Limitations and Failure Cases Although our SMR can effectively reconstruct 3D mesh object on multiple category-specific datasets with only 2D silhouette annotations, there are still two main limitations to be overcome in the near future: Limitation 1: Silhouette Annotations. It still requires silhouette annotations. For the sake of simplification, we have not considered the influence of background information. For some category-specific objects, such as horse, we annotate their silhouette annotations by detectron2. Thus the incorrect annotations will influence the reconstructed results. The failure case is illustrated in Fig. 2. Limitation 2: Non-Rigid Objects. It is not fit for non-grid objects, like the humans or flowers. Since it is difficult to determine the canonical viewpoints and the topological limitations of mesh representation. To avoid these limitations, we will further predict the silhouette masks by self-supervised learning and build a fully unsupervised 3D reconstruction model for more general objects. 3. More Reconstruction Results 3D Reconstruction On ShapeNet On the Shapenet dataset, since our SMR aims to model category-specific object, we perform experiments on these 13 categories one-byone. We introduce the ground truth camera parameters so as to evaluate the reconstructed accuracy and compare with other supervised methods to demonstrate SMR’s effective5x5 Conv, 32 5x5 Conv, 64 5x5 Conv, 128