An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition

An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
复制标题

用于基于图像的序列识别的端到端可训练神经网络及其在场景文本识别中的应用

DOI:
10.1109/tpami.2016.2646371
复制
发表时间:
2017-11-01
影响因子:
23.6
通讯作者:
Yao, Cong
Yao, Cong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shi, Baoguang;Bai, Xiang;Yao, Cong

文献摘要

被引文献

相似文献

基于图像的序列识别是计算机视觉领域的一个长期研究课题。场景文本识别是基于图像序列识别中最重要和最具挑战性的任务之一。提出了一种新的神经网络结构,它集成了特征提取,序列建模和转录到一个统一的框架。与以前的场景文本识别系统相比,该架构具有四个独特的属性:(1)它是端到端的可训练的,与大多数现有的算法,其组件是单独训练和调整。(2)它自然地处理任意长度的序列,不涉及字符分割或水平尺度规范化。(3)它不受任何预定义词典的限制,在无词典和基于词典的场景文本识别任务中都取得了显著的成绩。(4)它生成了一个有效但更小的模型,这对现实世界的应用场景更实用。在IIIT-5 K、街景文本和ICDAR数据集上的实验表明,该算法优于现有算法。此外,该算法在基于图像的乐谱识别任务中表现良好,这明显验证了它的通用性。
Image-based sequence recognition has been a long-standing research topic in computer vision. In this paper, we investigate the problem of scene text recognition, which is among the most important and challenging tasks in image-based sequence recognition. A novel neural network architecture, which integrates feature extraction, sequence modeling and transcription into a unified framework, is proposed. Compared with previous systems for scene text recognition, the proposed architecture possesses four distinctive properties: (1) It is end-to-end trainable, in contrast to most of the existing algorithms whose components are separately trained and tuned. (2) It naturally handles sequences in arbitrary lengths, involving no character segmentation or horizontal scale normalization. (3) It is not confined to any predefined lexicon and achieves remarkable performances in both lexicon-free and lexicon-based scene text recognition tasks. (4) It generates an effective yet much smaller model, which is more practical for real-world application scenarios. The experiments on standard benchmarks, including the IIIT-5K, Street View Text and ICDAR datasets, demonstrate the superiority of the proposed algorithm over the prior arts. Moreover, the proposed algorithm performs well in the task of image-based music score recognition, which evidently verifies the generality of it.