Deep Regionlets for Object Detection

Deep Regionlets for Object Detection
复制标题

DOI:
10.1007/978-3-030-01252-6_49
复制
发表时间:
2017-12
期刊:
--
影响因子:
--
通讯作者:
Hongyu Xu;Xutao Lv;Xiaoyu Wang;Zhou Ren;R. Chellappa
Hongyu Xu;Xutao Lv;Xiaoyu Wang;Zhou Ren;R. Chellappa
中科院分区:
其他
文献类型:
--
作者:
Hongyu Xu;Xutao Lv;Xiaoyu Wang;Zhou Ren;R. Chellappa

文献摘要

被引文献

相似文献

本文通过在深度神经网络和传统检测模式之间建立桥梁,提出了一种新的目标检测框架“Deep Regionlets”,以实现准确的通用目标检测。基于区域集建模对象变形和多长宽比的能力,我们将区域集整合到一个端到端可训练的深度学习框架中。该框架由区域选择网络和深度区域学习模块组成。具体来说,给定一个检测边界框建议,区域选择网络提供了从哪里选择区域来学习特征的指导。区域学习模块侧重于局部特征的选择和转换,以减轻局部变化。为此,我们首先在检测框架内实现非矩形区域的选择,以适应物体外观的变化。此外,我们在区域集学习模块内设计了一个“门控网络”,以实现区域集的软选择和池化。Deep Regionlets框架是端到端的训练,无需额外的努力。我们在PASCAL VOC和Microsoft COCO数据集上进行消融研究并进行广泛的实验。即使没有额外的分割标签,所提出的框架也优于最先进的算法,如RetinaNet和Mask R-CNN。
In this paper, we propose a novel object detection framework named" Deep Regionlets" by establishing a bridge between deep neural networks and conventional detection schema for accurate generic object detection. Motivated by the abilities of regionlets for modeling object deformation and multiple aspect ratios, we incorporate regionlets into an end-to-end trainable deep learning framework. The deep regionlets framework consists of a region selection network and a deep regionlet learning module. Specifically, given a detection bounding box proposal, the region selection network provides guidance on where to select regions to learn the features from. The regionlet learning module focuses on local feature selection and transformation to alleviate local variations. To this end, we first realize non-rectangular region selection within the detection framework to accommodate variations in object appearance. Moreover, we design a``gating network" within the regionlet leaning module to enable soft regionlet selection and pooling. The Deep Regionlets framework is trained end-to-end without additional efforts. We perform ablation studies and conduct extensive experiments on the PASCAL VOC and Microsoft COCO datasets. The proposed framework outperforms state-of-the-art algorithms, such as RetinaNet and Mask R-CNN, even without additional segmentation labels.