Object based Scene Representations using Fisher Scores of Local Subspace Projections

Object based Scene Representations using Fisher Scores of Local Subspace Projections
复制标题

DOI:
--
复制
发表时间:
2016-12
期刊:
--
影响因子:
--
通讯作者:
Mandar Dixit;N. Vasconcelos
Mandar Dixit;N. Vasconcelos
中科院分区:
其他
文献类型:
--
作者:
Mandar Dixit;N. Vasconcelos

文献摘要

被引文献

相似文献

一些工作已经表明,深度CNN可以很容易地在数据集之间传输,例如从ImageNet上的对象识别到Pascal VOC上的对象检测。然而,不太清楚的是CNN跨任务传输知识的能力。这种转移的一个常见例子是场景分类的问题,它应该利用局部对象检测来识别整体视觉概念。虽然这个问题目前是用Fisher向量表示来解决的,但现在这些表示对于现代CNN提取的高维和高度非线性特征是无效的。有人认为,这主要是由于依赖于一个模型,即对角协方差的高斯混合,它捕获CNN特征的二阶统计量的能力非常有限。这个问题是通过采用一个更好的模型,混合因子分析器(MFA),它近似的非线性数据流形的局部子空间的集合来解决。推导出相对于MFA的Fisher分数(MFA-FS),并提出作为整体图像分类器的图像表示。大量的实验表明,MFA-FS具有对象到场景传输的最先进性能,并且这种传输实际上优于从大型场景数据集训练场景CNN。这两种表示也被证明是互补的,在这个意义上,他们的组合优于每一个表示本身。组合后,它们会产生最先进的场景分类器。
Several works have shown that deep CNNs can be easily transferred across datasets, e.g. the transfer from object recognition on ImageNet to object detection on Pascal VOC. Less clear, however, is the ability of CNNs to transfer knowledge across tasks. A common example of such transfer is the problem of scene classification, that should leverage localized object detections to recognize holistic visual concepts. While this problems is currently addressed with Fisher vector representations, these are now shown ineffective for the high-dimensional and highly non-linear features extracted by modern CNNs. It is argued that this is mostly due to the reliance on a model, the Gaussian mixture of diagonal covariances, which has a very limited ability to capture the second order statistics of CNN features. This problem is addressed by the adoption of a better model, the mixture of factor analyzers (MFA), which approximates the non-linear data manifold by a collection of local sub-spaces. The Fisher score with respect to the MFA (MFA-FS) is derived and proposed as an image representation for holistic image classifiers. Extensive experiments show that the MFA-FS has state of the art performance for object-to-scene transfer and this transfer actually outperforms the training of a scene CNN from a large scene dataset. The two representations are also shown to be complementary, in the sense that their combination outperforms each of the representations by itself. When combined, they produce a state-of-the-art scene classifier.