Region-Based Convolutional Networks for Accurate Object Detection and Segmentation

Region-Based Convolutional Networks for Accurate Object Detection and Segmentation
复制标题

DOI:
10.1109/tpami.2015.2437384
复制
发表时间:
2016-01-01
影响因子:
23.6
通讯作者:
Malik, Jitendra
Malik, Jitendra
中科院分区:
计算机科学1区
文献类型:
--
作者:
Girshick, Ross;Donahue, Jeff;Malik, Jitendra

文献摘要

被引文献

相似文献

根据标准的帕斯卡VOC挑战赛数据集衡量的目标检测性能,在比赛的最后几年停滞不前。表现最好的方法是复杂的集成系统,通常将多个低级别图像特征与高级背景相结合。在本文中,我们提出了一种简单且可扩展的检测算法,与之前在VOC 2012上的最佳结果相比,平均平均精度(MAP)提高了50%以上-实现了62.4%的MAP。我们的方法结合了两个想法:(1)可以将高容量卷积网络(CNN)应用于自下而上的区域建议,以定位和分割对象;(2)当标记的训练数据稀缺时,辅助任务的有监督预训练,然后是特定领域的微调,显著提高了性能。由于我们将区域建议与CNN相结合,因此我们将所得到的模型称为R-CNN或基于区域的卷积网络。整个系统的源代码可在http://www.cs.berkeley.edu/similar to rbg/rcnn上获得。
Object detection performance, as measured on the canonical PASCAL VOC Challenge datasets, plateaued in the final years of the competition. The best-performing methods were complex ensemble systems that typically combined multiple low-level image features with high-level context. In this paper, we propose a simple and scalable detection algorithm that improves mean average precision (mAP) by more than 50 percent relative to the previous best result on VOC 2012-achieving a mAP of 62.4 percent. Our approach combines two ideas: (1) one can apply high-capacity convolutional networks (CNNs) to bottom-up region proposals in order to localize and segment objects and (2) when labeled training data are scarce, supervised pre-training for an auxiliary task, followed by domain-specific fine-tuning, boosts performance significantly. Since we combine region proposals with CNNs, we call the resulting model an R-CNN or Region-based Convolutional Network. Source code for the complete system is available at http://www.cs.berkeley.edu/similar to rbg/rcnn.