An Improved Scene Text Extraction Method Using Conditional Random Field and Optical Character Recognition
An Improved Scene Text Extraction Method Using Conditional Random Field and Optical Character Recognition
复制标题
DOI:
10.1109/icdar.2011.148
复制
发表时间:
2011-09
期刊:
影响因子:
--
通讯作者:
Hongwei Zhang;Changsong Liu;Cheng Yang;Xiaoqing Ding;Kongqiao Wang
中科院分区:
文献类型:
--
作者:
Hongwei Zhang;Changsong Liu;Cheng Yang;Xiaoqing Ding;Kongqiao Wang
Over the past few years, research on scene text extraction has developed rapidly. Recently, condition random field (CRF) has been used to give connected components (CCs) 'text' or 'non-text' labels. However, a burning issue in CRF model comes from multiple text lines extraction. In this paper, we propose a two-step iterative CRF algorithm with a Belief Propagation inference and an OCR filtering stage. Two kinds of neighborhood relationship graph are used in the respective iterations for extracting multiple text lines. Furthermore, OCR confidence is used as an indicator for identifying the text regions, while a traditional OCR filter module only considered the recognition results. The first CRF iteration aims at finding certain text CCs, especially in multiple text lines, and sending uncertain CCs to the second iteration. The second iteration gives second chance for the uncertain CCs and filter false alarm CCs with the help of OCR. Experiments based on the public dataset of ICDAR 2005 prove that the proposed method is comparative with the existing algorithms.