DeepText: A Unified Framework for Text Proposal Generation and Text Detection in Natural Images

DeepText: A Unified Framework for Text Proposal Generation and Text Detection in Natural Images
复制标题

DOI:
--
复制
发表时间:
2016-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhuoyao Zhong;Lianwen Jin;Shuye Zhang;Ziyong Feng
Zhuoyao Zhong;Lianwen Jin;Shuye Zhang;Ziyong Feng
中科院分区:
其他
文献类型:
--
作者:
Zhuoyao Zhong;Lianwen Jin;Shuye Zhang;Ziyong Feng

文献摘要

被引文献

相似文献

本文提出了一种基于完全卷积神经网络(CNN)的自然图像文本区域建议生成和文本检测的统一框架DeepText。首先,我们提出了初始区域提案网络,并设计了一组文本特征优先边界框来实现只有100个候选提案的高词汇召回率。接下来,我们提出了一个强大的文本检测网络,它嵌入了歧义文本类别(ATC)信息和多层感兴趣区域池(MLRP),用于文本和非文本分类和准确定位。最后,我们应用迭代包围盒投票方案,以互补的方式追求高召回率,并引入过滤算法来保留最合适的包围盒,同时去除每个文本实例的冗余内框和外框。在ICDAR 2011和2013年的稳健文本检测基准上,我们的方法获得了0.83和0.85的F度量,超过了之前的最先进结果。
In this paper, we develop a novel unified framework called DeepText for text region proposal generation and text detection in natural images via a fully convolutional neural network (CNN). First, we propose the inception region proposal network (Inception-RPN) and design a set of text characteristic prior bounding boxes to achieve high word recall with only hundred level candidate proposals. Next, we present a powerful textdetection network that embeds ambiguous text category (ATC) information and multilevel region-of-interest pooling (MLRP) for text and non-text classification and accurate localization. Finally, we apply an iterative bounding box voting scheme to pursue high recall in a complementary manner and introduce a filtering algorithm to retain the most suitable bounding box, while removing redundant inner and outer boxes for each text instance. Our approach achieves an F-measure of 0.83 and 0.85 on the ICDAR 2011 and 2013 robust text detection benchmarks, outperforming previous state-of-the-art results.