Multi-oriented Text Detection with Fully Convolutional Networks

Multi-oriented Text Detection with Fully Convolutional Networks
复制标题

DOI:
10.1109/cvpr.2016.451
复制
发表时间:
2016-04
期刊:
2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Zheng Zhang;Chengquan Zhang;Wei Shen;C. Yao;Wenyu Liu-;X. Bai
Zheng Zhang;Chengquan Zhang;Wei Shen;C. Yao;Wenyu Liu-;X. Bai
中科院分区:
其他
文献类型:
--
作者:
Zheng Zhang;Chengquan Zhang;Wei Shen;C. Yao;Wenyu Liu-;X. Bai

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的方法在自然图像中的文本检测。本地和全球的线索都考虑到本地化的文本行在一个由粗到细的过程。首先,训练全卷积网络(FCN)模型,以整体方式预测文本区域的显著图。然后,通过结合显著图和字符分量来估计文本行假设。最后,另一个FCN分类器被用来预测每个字符的质心,以消除错误的假设。该框架一般用于处理多种方向、语言和字体的文本。所提出的方法在三个文本检测基准:MSRA-TD 500,ICDAR 2015和ICDAR 2013上始终达到最先进的性能。
In this paper, we propose a novel approach for text detection in natural images. Both local and global cues are taken into account for localizing text lines in a coarse-to-fine procedure. First, a Fully Convolutional Network (FCN) model is trained to predict the salient map of text regions in a holistic manner. Then, text line hypotheses are estimated by combining the salient map and character components. Finally, another FCN classifier is used to predict the centroid of each character, in order to remove the false hypotheses. The framework is general for handling text in multiple orientations, languages and fonts. The proposed method consistently achieves the state-of-the-art performance on three text detection benchmarks: MSRA-TD500, ICDAR2015 and ICDAR2013.