HCP: A Flexible CNN Framework for Multi-Label Image Classification

HCP: A Flexible CNN Framework for Multi-Label Image Classification
复制标题

HCP:用于多标签图像分类的灵活 CNN 框架

DOI:
10.1109/tpami.2015.2491929
复制
发表时间:
2016-09-01
影响因子:
23.6
通讯作者:
Yan, Shuicheng
Yan, Shuicheng
中科院分区:
计算机科学1区
文献类型:
--
作者:
Wei, Yunchao;Xia, Wei;Yan, Shuicheng

文献摘要

被引文献

相似文献

卷积神经网络(CNN)在单标签图像分类任务中表现出了有希望的性能。但是,CNN如何最好地使用多标签图像仍然是一个空旷的问题,这主要是由于复杂的基础对象布局和多标签训练图像不足。在这项工作中,我们提出了一个灵活的深CNN基础架构,称为假设-CNN-Pooling(HCP),其中任意数量的对象段假设作为输入,然后将共享的CNN与每个假设相连,最后是CNN来自不同假设的输出结果与最大池汇总在一起,以产生最终的多标签预测。这种灵活的深CNN基础架构的一些独特特征包括:1)训练不需要地面真相边界框信息; 2)整个HCP基础设施可能对可能的嘈杂和/或冗余假设具有鲁棒性; 3)共享的CNN是灵活的,可以通过大型单标签图像数据集进行精心训练,例如Imagenet; 4)它可以自然输出多标签预测结果。 Pascal VOC 2007和VOC 2012多标签图像数据集的实验结果很好地证明了拟议的HCP基础架构优于其他最先进的机构。特别是,基于VOC 2012数据集中的手工制作的功能,该地图仅由HCP乘HCP达到90.5%,在与我们的互补结果融合后的93.2%。
Convolutional Neural Network (CNN) has demonstrated promising performance in single-label image classification tasks. However, how CNN best copes with multi-label images still remains an open problem, mainly due to the complex underlying object layouts and insufficient multi-label training images. In this work, we propose a flexible deep CNN infrastructure, called Hypotheses-CNN-Pooling (HCP), where an arbitrary number of object segment hypotheses are taken as the inputs, then a shared CNN is connected with each hypothesis, and finally the CNN output results from different hypotheses are aggregated with max pooling to produce the ultimate multi-label predictions. Some unique characteristics of this flexible deep CNN infrastructure include: 1) no ground-truth bounding box information is required for training; 2) the whole HCP infrastructure is robust to possibly noisy and/or redundant hypotheses; 3) the shared CNN is flexible and can be well pre-trained with a large-scale single-label image dataset, e.g., ImageNet; and 4) it may naturally output multi-label prediction results. Experimental results on Pascal VOC 2007 and VOC 2012 multi-label image datasets well demonstrate the superiority of the proposed HCP infrastructure over other state-of-the-arts. In particular, the mAP reaches 90.5% by HCP only and 93.2% after the fusion with our complementary result in [12] based on hand-crafted features on the VOC 2012 dataset.