Crafting GBD-Net for Object Detection

Crafting GBD-Net for Object Detection
复制标题

DOI:
10.1109/tpami.2017.2745563
复制
发表时间:
2016-10
影响因子:
23.6
通讯作者:
Xingyu Zeng;Wanli Ouyang;Junjie Yan;Hongsheng Li;Tong Xiao;Kun Wang;Yu Liu;Yucong Zhou;Binh Yang;Zhe Wang;Hui Zhou;Xiaogang Wang
Xingyu Zeng;Wanli Ouyang;Junjie Yan;Hongsheng Li;Tong Xiao;Kun Wang;Yu Liu;Yucong Zhou;Binh Yang;Zhe Wang;Hui Zhou;Xiaogang Wang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xingyu Zeng;Wanli Ouyang;Junjie Yan;Hongsheng Li;Tong Xiao;Kun Wang;Yu Liu;Yucong Zhou;Binh Yang;Zhe Wang;Hui Zhou;Xiaogang Wang

文献摘要

被引文献

相似文献

在目标检测中,来自不同大小和分辨率的多个支持区域的视觉线索在对候选框进行分类时是互补的。有效地整合这些区域的局部和上下文视觉线索已成为目标检测中的一个基本问题。在本文中,我们提出了一种门控双向CNN(GBD-Net),用于在特征学习和特征提取过程中在来自不同支持区域的特征之间传递消息。这种消息传递可以通过在两个方向上的相邻支持区域之间的卷积来实现,并且可以在各个层中进行。因此,局部和上下文视觉模式可以通过学习它们的非线性关系来验证彼此的存在,并且它们的紧密交互以更复杂的方式建模。它还表明,消息传递并不总是有帮助的,但依赖于个别样本。因此,需要门控函数来控制消息传输,其开或关由来自输入样本的额外视觉证据控制。通过在ImageNet、Pascal VOC 2007和Microsoft COCO三个目标检测数据集上的实验,验证了GBD-Net的有效性。除了GBD-Net,本文还展示了我们赢得2016年ImageNet对象检测挑战赛的方法的细节,源代码可在https://github.com/craftGBD/craftGBD上获得。在该系统中,改进的GBD网络,新的预训练方案和更好的区域建议设计。我们还展示了不同的网络结构和现有的对象检测技术的有效性,如多尺度测试,左右翻转,边界框投票,NMS和上下文。
The visual cues from multiple support regions of different sizes and resolutions are complementary in classifying a candidate box in object detection. Effective integration of local and contextual visual cues from these regions has become a fundamental problem in object detection. In this paper, we propose a gated bi-directional CNN (GBD-Net) to pass messages among features from different support regions during both feature learning and feature extraction. Such message passing can be implemented through convolution between neighboring support regions in two directions and can be conducted in various layers. Therefore, local and contextual visual patterns can validate the existence of each other by learning their nonlinear relationships and their close interactions are modeled in a more complex way. It is also shown that message passing is not always helpful but dependent on individual samples. Gated functions are therefore needed to control message transmission, whose on-or-offs are controlled by extra visual evidence from the input sample. The effectiveness of GBD-Net is shown through experiments on three object detection datasets, ImageNet, Pascal VOC2007 and Microsoft COCO. Besides the GBD-Net, this paper also shows the details of our approach in winning the ImageNet object detection challenge of 2016, with source code provided on https://github.com/craftGBD/craftGBD. In this winning system, the modified GBD-Net, new pretraining scheme and better region proposal designs are provided. We also show the effectiveness of different network structures and existing techniques for object detection, such as multi-scale testing, left-right flip, bounding box voting, NMS, and context.