Attentive Contexts for Object Detection

Attentive Contexts for Object Detection
复制标题

DOI:
10.1109/tmm.2016.2642789
复制
发表时间:
2017-05-01
影响因子:
7.3
通讯作者:
Yan, Shuicheng
Yan, Shuicheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Jianan;Wei, Yunchao;Yan, Shuicheng

文献摘要

被引文献

相似文献

现代基于深度神经网络的目标检测方法通常使用候选提案的内部特征对其进行分类。然而,被认为对目标检测有价值的全局和局部周围环境尚未被现有方法充分利用。在这项工作中,我们朝着理解什么是提取和利用上下文信息以促进实践中目标检测的稳健实践迈出了一步。具体来说,我们考虑了以下两个问题:“如何识别有用的全局上下文信息来检测某个对象?”以及“如何利用提案周围的局部上下文来更好地推断其内容?”我们通过开发一种新的基于关注上下文卷积神经网络(AC-CNN)的目标检测模型,为这些问题提供了初步的答案。AC-CNN有效地将全局和局部上下文信息融合到基于区域的CNN(如快速R-CNN和更快R-CNN)检测框架中,并提供更好的目标检测性能。它由一个基于注意的全局语境化(AGC)子网和一个多尺度局部语境化(MLC)子网组成。为了捕捉全局上下文,AGC子网通过多个堆叠的长短期记忆层,为输入图像循环生成一个注意地图,以突出显示有用的全局上下文位置。为了捕获周围的本地上下文,MLC子网在多个尺度上利用每个特定提案的内部和外部上下文信息。然后将全局和局部上下文融合在一起,以做出检测的最终决策。在PASCAL VOC 2007和VOC 2012上进行的大量实验很好地证明了所提出的AC-CNN优于已建立的基线。
Modern deep neural network-based object detection methods typically classify candidate proposals using their interior features. However, global and local surrounding contexts that are believed to be valuable for object detection are not fully exploited by existing methods yet. In this work, we take a step towards understanding what is a robust practice to extract and utilize contextual information to facilitate object detection in practice. Specifically, we consider the following two questions: "how to identify useful global contextual information for detecting a certain object?" and "how to exploit local context surrounding a proposal for better inferring its contents?" We provide preliminary answers to these questions through developing a novel attention to context convolution neural network (AC-CNN)-based object detection model. AC-CNN effectively incorporates global and local contextual information into the region-based CNN (e.g., fast R-CNN and faster R-CNN) detection framework and provides better object detection performance. It consists of one attention-based global contextualized (AGC) subnetwork and one multi-scale local contextualized (MLC) subnetwork. To capture global context, the AGC subnetwork recurrently generates an attention map for an input image to highlight useful global contextual locations, through multiple stacked long short-term memory layers. For capturing surrounding local context, the MLC subnetwork exploits both the inside and outside contextual information of each specific proposal at multiple scales. The global and local context are then fused together for making the final decision for detection. Extensive experiments on PASCAL VOC 2007 and VOC 2012 well demonstrate the superiority of the proposed AC-CNN over well-established baselines.