Object Pose Estimation using Mid-level Visual Representations

Object Pose Estimation using Mid-level Visual Representations
复制标题

DOI:
10.1109/iros47612.2022.9981452
复制
发表时间:
2022-03
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Negar Nejatishahidin;Pooya Fayyazsanavi;J. Kosecka
Negar Nejatishahidin;Pooya Fayyazsanavi;J. Kosecka
中科院分区:
其他
文献类型:
--
作者:
Negar Nejatishahidin;Pooya Fayyazsanavi;J. Kosecka

文献摘要

被引文献

相似文献

该工作提出了一种新的目标类别姿态估计模型,该模型可以有效地转移到先前未知的环境中。用于姿态估计的深卷积网络模型(CNN)通常在专门为目标检测、姿态估计或3D重建而策划的数据集上进行训练和评估,这需要大量的训练数据。在这项工作中,我们提出了一个用于姿态估计的模型,该模型可以用少量的数据进行训练,并且建立在通用的中层表示[33]的基础上(例如,表面法线估计和重着色)。这些表示是在大数据集上训练的,不需要姿势和对象注释。随后,通过一个利用物体掩码和轮廓检索的小型CNN神经网络来改进预测。该方法在Pix3D数据集[26]上取得了优异的性能,并且在仅有25%的训练数据可用时,比现有模型提高了近35%。我们表明,当涉及到泛化和转移到新环境时,该方法是有利的。为此,我们在挑战活动视觉数据集[1]上为常见的家具类别引入了一种新的姿势估计基准,并对在Pix3D数据集上训练的模型进行了评估。
This work proposes a novel pose estimation model for object categories that can be effectively transferred to pre-viously unseen environments. The deep convolutional network models (CNN) for pose estimation are typically trained and evaluated on datasets specifically curated for object detection, pose estimation, or 3D reconstruction, which requires large amounts of training data. In this work, we propose a model for pose estimation that can be trained with small amount of data and is built on the top of generic mid-level represen-tations [33] (e.g. surface normal estimation and re-shading). These representations are trained on a large dataset without requiring pose and object annotations. Later on, the predictions are refined with a small CNN neural network that exploits object masks and silhouette retrieval. The presented approach achieves superior performance on the Pix3D dataset [26] and shows nearly 35 % improvement over the existing models when only 25 % of the training data is available. We show that the approach is favorable when it comes to generalization and transfer to novel environments. Towards this end, we introduce a new pose estimation benchmark for commonly encountered furniture categories on challenging Active Vision Dataset [1] and evaluated the models trained on the Pix3D dataset.