Reading Scene Text in Deep Convolutional Sequences

Reading Scene Text in Deep Convolutional Sequences
复制标题

DOI:
10.1609/aaai.v30i1.10465
复制
发表时间:
2015-06
期刊:
--
影响因子:
--
通讯作者:
Pan He;Weilin Huang;Y. Qiao;Chen Change Loy;Xiaoou Tang
Pan He;Weilin Huang;Y. Qiao;Chen Change Loy;Xiaoou Tang
中科院分区:
其他
文献类型:
--
作者:
Pan He;Weilin Huang;Y. Qiao;Chen Change Loy;Xiaoou Tang

文献摘要

被引文献

相似文献

我们开发了一个深度文本递归网络(DTRN),将场景文本阅读视为序列标记问题。我们利用深度卷积神经网络的最新进展,从整个单词图像中生成有序的高级序列,避免了困难的字符分割问题。然后,开发了一个基于长短期记忆(LSTM)的深度递归模型,以鲁棒地识别生成的CNN序列,这与大多数现有的独立识别每个字符的方法不同。与现有的场景文本识别方法相比,我们的模型具有许多吸引人的特性:(i)它可以通过利用有意义的上下文信息来识别高度模糊的单词,使其能够在没有预处理或后处理的情况下可靠地工作;(ii)深度CNN特征对各种图像失真具有鲁棒性;(3)它保留了词意象中的显式顺序信息,这是区分词串所必需的;(iv)该模型不依赖于预定义的字典,并且可以处理未知词和任意字符串。它在几个基准上取得了令人印象深刻的结果,大大推进了最先进的水平。
We develop a Deep-Text Recurrent Network (DTRN)that regards scene text reading as a sequence labelling problem. We leverage recent advances of deep convolutional neural networks to generate an ordered highlevel sequence from a whole word image, avoiding the difficult character segmentation problem. Then a deep recurrent model, building on long short-term memory (LSTM), is developed to robustly recognize the generated CNN sequences, departing from most existing approaches recognising each character independently. Our model has a number of appealing properties in comparison to existing scene text recognition methods: (i) It can recognise highly ambiguous words by leveraging meaningful context information, allowing it to work reliably without either pre- or post-processing; (ii) the deep CNN feature is robust to various image distortions; (iii) it retains the explicit order information in word image, which is essential to discriminate word strings; (iv) the model does not depend on pre-defined dictionary, and it can process unknown words and arbitrary strings. It achieves impressive results on several benchmarks, advancing the-state-of-the-art substantially.