HCR-Net: A Hybrid of Classification and Regression Network for Object Pose Estimation

HCR-Net: A Hybrid of Classification and Regression Network for Object Pose Estimation
复制标题

DOI:
10.24963/ijcai.2018/141
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Zairan Wang;Weiming Li;Yueying Kao;Dongqing Zou;Qiang Wang;Minsu Ahn;Sunghoon Hong
Zairan Wang;Weiming Li;Yueying Kao;Dongqing Zou;Qiang Wang;Minsu Ahn;Sunghoon Hong
中科院分区:
其他
文献类型:
--
作者:
Zairan Wang;Weiming Li;Yueying Kao;Dongqing Zou;Qiang Wang;Minsu Ahn;Sunghoon Hong

文献摘要

被引文献

相似文献

从一幅图像中估计物体的姿态是计算机视觉和机器人技术中的一个基本和具有挑战性的问题。通常,当前的方法将姿态估计视为分类或回归问题。然而,基于回归的方法通常遭受训练数据不平衡的问题,而分类方法难以区分附近的姿势。本文提出了一种混合CNN模型,我们称之为HCR-Net,它集成了分类网络和回归网络,以处理这些问题。我们的模型的灵感来自于回归方法可以在均匀分布的数据集上获得更好的准确性,而分类方法对于粗略量化的姿势更有效,即使数据集不平衡。分类方法和回归方法本质上是相辅相成的。因此,我们将它们以混合方式集成到神经网络中,并使用两个新的损失函数进行端到端的训练。因此,我们的方法超越了最先进的方法,即使是不平衡的训练数据和更少的数据增强。在具有挑战性的Pascal 3D+数据库上的实验结果表明,我们的方法明显优于最先进的方法,ACC和AVP指标分别提高了4%和6%。
Object pose estimation from a single image is a fundamental and challenging problem in computer vision and robotics. Generally, current methods treat pose estimation as a classification or a regression problem. However, regression based methods usually suffer from the issue of imbalanced training data, while classification methods are difficult to discriminate nearby poses. In this paper, a hybrid CNN model, which we call it HCR-Net that integrates both a classification network and a regression network, is proposed to deal with these issues. Our model is inspired by that regression methods can get better accuracy on homogeneously distributed datasets while classification methods are more effective for coarse quantization of the poses even if the dataset is not well balanced. The classification methods and the regression methods essentially complement each other. Thus we integrate both them into a neural network in a hybrid fashion and train it end-to-end with two novel loss functions. As a result, our method surpass the state-of-the-art methods, even with imbalanced training data and much less data augmentation. The experimental results on the challenging Pascal3D+ database demonstrate that our method outperforms the state-of-the-arts significantly, achieving improvements on ACC and AVP metrics up to 4% and 6%, respectively.