Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction

Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category Reconstruction
复制标题

DOI:
10.1109/iccv48922.2021.01072
复制
发表时间:
2021-09
期刊:
2021 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Jeremy Reizenstein;Roman Shapovalov;P. Henzler;L. Sbordone;Patrick Labatut;David Novotný
Jeremy Reizenstein;Roman Shapovalov;P. Henzler;L. Sbordone;Patrick Labatut;David Novotný
中科院分区:
其他
文献类型:
--
作者:
Jeremy Reizenstein;Roman Shapovalov;P. Henzler;L. Sbordone;Patrick Labatut;David Novotný

文献摘要

被引文献

相似文献

传统的学习3D对象类别的方法主要是在合成数据集上训练和评估的,这是因为无法获得真实的3D注释的以类别为中心的数据。我们的主要目标是通过收集与现有合成数据类似的真实世界数据来促进这一领域的进步。因此,这项工作的主要贡献是一个大型数据集,称为3D中的常见对象,其中包含用相机姿势和地面真实3D点云注释的对象类别的真实多视图图像。该数据集包含来自50个MS-COCO类别的近19,000个对象的近19,000个视频的总共150万帧,因此,它在类别和对象的数量上都明显大于备选方法。我们利用这个新的数据集对几种新的视图合成和以类别为中心的3D重建方法进行了首次大规模的“野外”评估。最后,我们贡献了NerFormer-一种新的神经渲染方法,它利用强大的Transformer来重建给定少量视图的对象。
Traditional approaches for learning 3D object categories have been predominantly trained and evaluated on synthetic datasets due to the unavailability of real 3D-annotated category-centric data. Our main goal is to facilitate advances in this field by collecting real-world data in a magnitude similar to the existing synthetic counterparts. The principal contribution of this work is thus a large-scale dataset, called Common Objects in 3D, with real multi-view images of object categories annotated with camera poses and ground truth 3D point clouds. The dataset contains a total of 1.5 million frames from nearly 19,000 videos capturing objects from 50 MS-COCO categories and, as such, it is significantly larger than alternatives both in terms of the number of categories and objects.We exploit this new dataset to conduct one of the first large-scale "in-the-wild" evaluations of several new-view-synthesis and category-centric 3D reconstruction methods. Finally, we contribute NerFormer - a novel neural rendering method that leverages the powerful Transformer to reconstruct an object given a small number of its views.