Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition

Convolutional recurrent neural networks with hidden Markov model bootstrap for scene text recognition
复制标题

用于场景文本识别的具有隐马尔可夫模型引导的卷积循环神经网络

DOI:
10.1049/iet-cvi.2016.0417
复制
发表时间:
2017-06
影响因子:
1.7
通讯作者:
Zhang Jun
Zhang Jun
中科院分区:
计算机科学4区
文献类型:
--
作者:
Wang Fenglei;Guo Qiang;Lei Jun;Zhang Jun

文献摘要

参考文献

相似文献

由于在不受约束的条件下外观高度可变,自然场景中的文本识别仍然是一个具有挑战性的问题。作者开发了一种系统,可以直接将场景文本图像转录为文本,无需进行字符分割。他们将问题表述为序列标记。他们使用深度卷积神经网络 (CNN) 来建模文本外观,并使用 RNN 来构建序列动力学,从而构建了卷积循环神经网络 (RNN)。这两个模型在建模能力上是互补的,因此集成在一起形成自由分割系统。他们训练高斯混合模型-隐马尔可夫模型来监督 CNN 模型的训练。该系统是数据驱动的,不需要手工标记的训练数据。他们的方法有几个吸引人的特性:(i)它可以识别任意长度的文本图像。 (ii) 识别过程不涉及复杂的字符分割。 (iii) 它是在仅具有单词级转录的场景文本图像上进行训练的。 (iv)它可以识别基于词典的文本或无词典的文本。所提出的系统在几个公共场景文本数据集(包括基于词典的和非词典的数据集)上实现了与现有技术水平的竞争性能比较。
Text recognition in natural scene remains a challenging problem due to the highly variable appearance in unconstrained condition. The authors develop a system that directly transcribes scene text images to text without character segmentation. They formulate the problem as sequence labelling. They build a convolutional recurrent neural network (RNN) by using deep convolutional neural networks (CNN) for modelling text appearance and RNNs for sequence dynamics. The two models are complementary in modelling capabilities and so integrated together to form the segmentation free system. They train a Gaussian mixture model-hidden Markov model to supervise the training of the CNN model. The system is data driven and needs no hand labelled training data. Their method has several appealing properties: (i) It can recognise arbitrary length text images. (ii) The recognition process does not involve sophisticated character segmentation. (iii) It is trained on scene text images with only word-level transcriptions. (iv) It can recognise both the lexicon-based or lexicon-free text. The proposed system achieves competitive performance comparison with the state of the art on several public scene text datasets, including both lexicon-based and non-lexicon ones.
DOI: 10.1109/cvpr.2016.451
发表时间: 2016-04
期刊: 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子: --
作者:
Zheng Zhang;Chengquan Zhang;Wei Shen;C. Yao;Wenyu Liu-;X. Bai
通讯作者: Zheng Zhang;Chengquan Zhang;Wei Shen;C. Yao;Wenyu Liu-;X. Bai
DOI: 10.1609/aaai.v30i1.10465
发表时间: 2015-06
期刊: --
影响因子: --
作者:
Pan He;Weilin Huang;Y. Qiao;Chen Change Loy;Xiaoou Tang
通讯作者: Pan He;Weilin Huang;Y. Qiao;Chen Change Loy;Xiaoou Tang
用于基于图像的序列识别的端到端可训练神经网络及其在场景文本识别中的应用
DOI: 10.1109/tpami.2016.2646371
发表时间: 2017-11-01
影响因子: 23.6
作者:
Shi, Baoguang;Bai, Xiang;Yao, Cong
通讯作者: Yao, Cong
DOI: 10.1109/cvpr.2014.515
发表时间: 2014-06
期刊: 2014 IEEE Conference on Computer Vision and Pattern Recognition
影响因子: --
作者:
C. Yao;X. Bai;Baoguang Shi;Wenyu Liu-
通讯作者: C. Yao;X. Bai;Baoguang Shi;Wenyu Liu-
DOI: 10.1109/tpami.2004.14
发表时间: 2004-06
影响因子: 23.6
作者:
A. Vinciarelli;Samy Bengio;H. Bunke
通讯作者: A. Vinciarelli;Samy Bengio;H. Bunke