Supplementary of Self-Supervised 3D Mesh Reconstruction from Single Images
Supplementary of Self-Supervised 3D Mesh Reconstruction from Single Images
复制标题
单图像自监督 3D 网格重建的补充
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Jiaya Jia
中科院分区:
文献类型:
--
作者:
T. Hu;Liwei Wang;Xiaogang Xu;Shu Liu;Jiaya Jia
1. Network Structure The input data X = [I,M ] ∈ RH×W×4 is concatenated by the original RGB single image I ∈ RH×W×3 and its silhouette mask M ∈ RH×W×1. Our 3D mesh reconstruction model consists of 4 sub-encoders, i.e. Camera Encoder Ec, Light Encoder El, Shape Encoder Es, and Texture Encoder Et. These sub-encoders separately encode corresponding attribute feature to avoid mutual effect, as illustrated in Fig. 1. The landmark feature extractor Ef is the conventional U-Net network whose output resolution is equal to the input resolution H ×W and output channel is 128. The landmark classification network is a MLP network with three layers whose output dimension is V . 2. Limitations and Failure Cases Although our SMR can effectively reconstruct 3D mesh object on multiple category-specific datasets with only 2D silhouette annotations, there are still two main limitations to be overcome in the near future: Limitation 1: Silhouette Annotations. It still requires silhouette annotations. For the sake of simplification, we have not considered the influence of background information. For some category-specific objects, such as horse, we annotate their silhouette annotations by detectron2. Thus the incorrect annotations will influence the reconstructed results. The failure case is illustrated in Fig. 2. Limitation 2: Non-Rigid Objects. It is not fit for non-grid objects, like the humans or flowers. Since it is difficult to determine the canonical viewpoints and the topological limitations of mesh representation. To avoid these limitations, we will further predict the silhouette masks by self-supervised learning and build a fully unsupervised 3D reconstruction model for more general objects. 3. More Reconstruction Results 3D Reconstruction On ShapeNet On the Shapenet dataset, since our SMR aims to model category-specific object, we perform experiments on these 13 categories one-byone. We introduce the ground truth camera parameters so as to evaluate the reconstructed accuracy and compare with other supervised methods to demonstrate SMR’s effective5x5 Conv, 32 5x5 Conv, 64 5x5 Conv, 128