A Hybrid Approach to Detect and Localize Texts in Natural Scene Images

A Hybrid Approach to Detect and Localize Texts in Natural Scene Images
复制标题

DOI:
10.1109/tip.2010.2070803
复制
发表时间:
2011-03-01
影响因子:
10.6
通讯作者:
Liu, Cheng-Lin
Liu, Cheng-Lin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Pan, Yi-Feng;Hou, Xinwen;Liu, Cheng-Lin

文献摘要

被引文献

相似文献

自然场景图像中的文本检测和定位是基于内容的图像分析的重要内容。由于背景复杂,光照不均匀,文本字体、大小和行方向的变化,该问题具有挑战性。在本文中,我们提出了一种混合的方法来鲁棒地检测和定位自然场景图像中的文本。设计了文本区域检测器,估计文本在图像金字塔中存在的置信度和尺度信息,通过局部二值化分割候选文本成分。为了有效地过滤掉非文本成分,提出了一种考虑一元成分属性和二元上下文成分关系的条件随机场(CRF)模型,该模型具有监督参数学习。最后,文本组件被分组为文本行/词与基于学习的能量最小化方法。由于所有三个阶段都是基于学习的,因此需要手动调整的参数很少。在ICDAR 2005竞赛数据集上的实验结果表明,与现有的方法相比,该方法具有更高的准确率和召回率。我们还在多语言图像数据集上评估了我们的方法,结果令人鼓舞。
Text detection and localization in natural scene images is important for content-based image analysis. This problem is challenging due to the complex background, the non-uniform illumination, the variations of text font, size and line orientation. In this paper, we present a hybrid approach to robustly detect and localize texts in natural scene images. A text region detector is designed to estimate the text existing confidence and scale information in image pyramid, which help segment candidate text components by local binarization. To efficiently filter out the non-text components, a conditional random field (CRF) model considering unary component properties and binary contextual component relationships with supervised parameter learning is proposed. Finally, text components are grouped into text lines/words with a learning-based energy minimization method. Since all the three stages are learning-based, there are very few parameters requiring manual tuning. Experimental results evaluated on the ICDAR 2005 competition dataset show that our approach yields higher precision and recall performance compared with state-of-the-art methods. We also evaluated our approach on a multilingual image dataset with promising results.