Scene text recognition using similarity and a lexicon with sparse belief propagation.

Scene text recognition using similarity and a lexicon with sparse belief propagation.
复制标题

使用相似性和具有稀疏置信传播的词典进行场景文本识别。

DOI:
10.1109/tpami.2009.38
复制
发表时间:
2009
影响因子:
23.6
通讯作者:
Hanson,AllenR
Hanson,AllenR
中科院分区:
计算机科学1区
文献类型:
--
作者:
Weinman,JerodJ;Learned-Miller,Erik;Hanson,AllenR

文献摘要

被引文献

相似文献

场景文本识别(STR)是对环境中任何地方的文本的识别,例如标志和店面。相对于文档识别,由于字体可变性、最小的语言上下文和不受控制的条件,它具有挑战性。许多可用于解决此问题的信息经常被忽略或顺序使用。人物形象之间的相似性往往被忽视,而这是有用的信息。由于语言先验,识别器可以为相同的字符分配不同的标签。直接相互比较字符,而不仅仅是比较模型,有助于确保相似的实例接收相同的标签。词汇可以提高识别的准确性,但通常是在事后使用。我们为STR引入了一个概率模型,该模型集成了相似性、语言属性和词法决策。通过稀疏信念传播加速推理,这是一种自下而上的方法,通过减少弱支持假设之间的依赖来缩短消息。通过将信息源融合在一个模型中,我们消除了由于顺序处理而导致的不可恢复的错误,提高了准确性。在实验结果中,从户外场景的标志图像中识别文本,结合相似性将字符识别错误率降低了19%,词典将单词识别错误率降低了35%,而稀疏信念传播在12倍加速的情况下将考虑的词典单词减少了99.9%,并且准确性没有损失。
Scene text recognition (STR) is the recognition of text anywhere in the environment, such as signs and storefronts. Relative to document recognition, it is challenging because of font variability, minimal language context, and uncontrolled conditions. Much information available to solve this problem is frequently ignored or used sequentially. Similarity between character images is often overlooked as useful information. Because of language priors, a recognizer may assign different labels to identical characters. Directly comparing characters to each other, rather than only a model, helps ensure that similar instances receive the same label. Lexicons improve recognition accuracy but are used post hoc. We introduce a probabilistic model for STR that integrates similarity, language properties, and lexical decision. Inference is accelerated with sparse belief propagation, a bottom-up method for shortening messages by reducing the dependency between weakly supported hypotheses. By fusing information sources in one model, we eliminate unrecoverable errors that result from sequential processing, improving accuracy. In experimental results recognizing text from images of signs in outdoor scenes, incorporating similarity reduces character recognition error by 19 percent, the lexicon reduces word recognition error by 35 percent, and sparse belief propagation reduces the lexicon words considered by 99.9 percent with a 12X speedup and no loss in accuracy.