Fisher vector for scene character recognition: A comprehensive evaluation

Fisher vector for scene character recognition: A comprehensive evaluation
复制标题

用于场景字符识别的 Fisher 向量:综合评价

DOI:
10.1016/j.patcog.2017.06.022
复制
发表时间:
2017-12
影响因子:
8
通讯作者:
Xiao Baihua
Xiao Baihua
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shi Cunzhao;Wang Yanna;Jia Fuxi;He Kun;Wang Chunheng;Xiao Baihua

文献摘要

参考文献

被引文献

相似文献

Fisher向量(FV)可以被看作是一个视觉词袋(BOW),它不仅对单词计数进行编码,还对高阶统计进行编码,与线性分类器配合使用,并在图像分类方面表现出良好的性能。对于字符识别,虽然标准的BOW已经被应用,但结果仍然不令人满意。本文将基于高斯混合模型(GMM)的视觉词汇的Fisher向量应用于字符识别,并结合了空间信息。我们在一系列具有挑战性的英文和数字字符识别数据集上,包括手写和场景字符识别数据集,给出了Fisher向量与线性分类器的综合评估。此外,我们还收集了两个中文场景字符识别数据集,以评估Fisher向量表示汉字的适用性。通过大量的实验,我们做了三个方面的工作:(1)证明了线性分类器的FV方法在字符识别中的性能优于大多数现有的字符识别方法,甚至优于基于CNN的方法,并且在训练样本不足的情况下,这种优势更加明显;(2)空间信息对汉字的表征是非常有用的,特别是对于结构复杂的汉字;(3)这一结果也暗示了FV具有代表新的未知类别的潜力,这是非常鼓舞人心的,因为对于大类别的中文场景字符来说,很难收集到足够的训练样本。
Fisher vector (FV), which could be seen as a bag of visual words (BOW) that encodes not only word counts but also higher-order statistics, works well with linear classifiers and has shown promising performance for image categorization. For character recognition, although standard BOW has been applied, the results are still not satisfactory. In this paper, we apply Fisher vector derived from Gaussian Mixture Models (GMM) based visual vocabularies on character recognition and integrate spatial information as well. We give a comprehensive evaluation of Fisher vector with linear classifier on a series of challenging English and digits character recognition datasets, including both the handwritten and scene character recognition ones. Moreover, we also collect two Chinese scene character recognition datasets to evaluate the suitability of Fisher vector to represent Chinese characters. Through extensive experiments we make three contributions: (1) we demonstrate that FV with linear classifier could outperform most of the state-of-the-art methods for character recognition, even the CNN based ones and the superiority is more obvious when training samples are insufficient to train the networks; (2) we show that additional spatial information is very useful for character representation, especially for Chinese ones, which have more complex structures; and (3) the results also imply the potential of FV to represent new unseen categories, which is quite inspiring since it is quite difficult to collect enough training samples for large-category Chinese scene characters.
DOI: 10.1023/b:visi.0000029664.99615.94
发表时间: 2004-11-01
影响因子: 19.5
作者:
Lowe, DG
通讯作者: Lowe, DG
DOI: --
发表时间: 2002
期刊: 2020 IEEE International Conference on Applied Superconductivity and Electromagnetic Devices (ASEMD)
影响因子: --
作者:
Gabriella Csurka;C. Dance;Lixin Fan;J. Willamowski;Cédric Bray
通讯作者: Gabriella Csurka;C. Dance;Lixin Fan;J. Willamowski;Cédric Bray
DOI: 10.1109/cvpr.2014.515
发表时间: 2014-06
期刊: 2014 IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者:
C. Yao;X. Bai;Baoguang Shi;Wenyu Liu-
通讯作者: C. Yao;X. Bai;Baoguang Shi;Wenyu Liu-
DOI: 10.1109/cvpr.2007.383266
发表时间: 2007-06
期刊: 2007 IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者:
Florent Perronnin;C. Dance
通讯作者: Florent Perronnin;C. Dance
DOI: 10.1109/icpr.2014.501
发表时间: 2014-08
期刊: 2014 22nd International Conference on Pattern Recognition
影响因子: --
作者:
Song Gao;Chunheng Wang;Baihua Xiao;Cunzhao Shi;Zhong Zhang
通讯作者: Song Gao;Chunheng Wang;Baihua Xiao;Cunzhao Shi;Zhong Zhang