A Weakly Supervised Fine Label Classifier Enhanced by Coarse Supervision

A Weakly Supervised Fine Label Classifier Enhanced by Coarse Supervision
复制标题

DOI:
10.1109/iccv.2019.00656
复制
发表时间:
2019-10
期刊:
2019 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Fariborz Taherkhani;Hadi Kazemi;Ali Dabouei;J. Dawson;N. Nasrabadi
Fariborz Taherkhani;Hadi Kazemi;Ali Dabouei;J. Dawson;N. Nasrabadi
中科院分区:
其他
文献类型:
--
作者:
Fariborz Taherkhani;Hadi Kazemi;Ali Dabouei;J. Dawson;N. Nasrabadi

文献摘要

被引文献

相似文献

对象通常以分层结构组织,其中每个粗略类别(例如,大猫)对应于若干细类别的超类(例如,猎豹、豹)。在相同的粗类别,但在不同的精细类别的对象分组,通常共享一组全球的视觉特征,但是,这些对象具有独特的本地属性,在一个精细的水平上表征他们。本文解决了在弱监督的方式,其中一个子集的图像被标记的精细标签,而其余的被标记的粗糙标签的精细图像分类的挑战。我们提出了一种新的深度模型,该模型利用粗糙图像来提高粗糙类别中精细图像的分类性能。我们的模型是一个端到端的框架,由卷积神经网络(CNN)组成,它使用精细和粗糙的图像来调整其参数。然后,CNN输出被扇出到两个单独的分支中,使得第一分支使用受监督的低秩自表达层将CNN输出投影到低秩子空间以捕获用于粗略分类的全局结构,而另一个分支使用受监督的稀疏自表达层将它们投影到稀疏子空间以捕获用于精细分类的局部结构。我们的深度模型使用粗糙图像与精细图像结合,通过在训练期间共享参数来共同探索低秩和稀疏子空间,这使得CNN获得的数据点被很好地投影到稀疏和低秩子空间进行分类。
Objects are usually organized in a hierarchical structure in which each coarse category (e.g., big cat) corresponds to a superclass of several fine categories (e.g., cheetah, leopard). The objects grouped within the same coarse category, but in different fine categories, usually share a set of global visual features; however, these objects have distinctive local properties that characterize them at a fine level. This paper addresses the challenge of fine image classification in a weakly supervised fashion, whereby a subset of images is tagged by fine labels, while the remaining are tagged by coarse labels. We propose a new deep model that leverages coarse images to improve the classification performance of fine images within the coarse category. Our model is an end to end framework consisting of a Convolutional Neural Network (CNN) which uses both fine and coarse images to tune its parameters. The CNN outputs are then fanned out into two separate branches such that the first branch uses a supervised low rank self expressive layer to project the CNN outputs to the low rank subspaces to capture the global structures for the coarse classification, while the other branch uses a supervised sparse self expressive layer to project them to the sparse subspaces to capture the local structures for the fine classification. Our deep model uses coarse images in conjunction with fine images to jointly explore the low rank and sparse subspaces by sharing the parameters during the training which causes the data points obtained by the CNN to be well-projected to both sparse and low rank subspaces for classification.