Detecting and reading text in natural scenes

Detecting and reading text in natural scenes
复制标题

DOI:
10.1109/cvpr.2004.77
复制
发表时间:
2004-06
期刊:
Proceedings of the 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2004. CVPR 2004.
影响因子:
--
通讯作者:
Xiangrong Chen;A. Yuille
Xiangrong Chen;A. Yuille
中科院分区:
其他
文献类型:
--
作者:
Xiangrong Chen;A. Yuille

文献摘要

被引文献

相似文献

本文给出了一种检测和读取自然图像中文本的算法。该算法旨在供盲人和视障人士在城市场景中行走时使用。我们首先获得盲人和视力正常的受试者拍摄的城市图像的数据集。从这个数据集中,我们手动标记并提取文本区域。接下来,我们对文本区域进行统计分析,以确定哪些图像特征是文本的可靠指标并且具有低熵(即所有文本图像的特征响应都是相似的)。我们通过使用文本上和文本外的特征响应的联合概率来获得弱分类器。这些弱分类器用作 AdaBoost 机器学习算法的输入来训练强分类器。在实践中,我们训练了包含 79 个特征的 4 个强分类器的级联。自适应二值化和扩展算法应用于级联分类器选择的那些区域。商业 OCR 软件用于读取文本或将其拒绝为非文本区域。整体算法在测试集上的成功率超过 90%(通过完整检测和阅读文本来评估),并且未读文本通常很小且距离观看者较远。
This paper gives an algorithm for detecting and reading text in natural images. The algorithm is intended for use by blind and visually impaired subjects walking through city scenes. We first obtain a dataset of city images taken by blind and normally sighted subjects. From this dataset, we manually label and extract the text regions. Next we perform statistical analysis of the text regions to determine which image features are reliable indicators of text and have low entropy (i.e. feature response is similar for all text images). We obtain weak classifiers by using joint probabilities for feature responses on and off text. These weak classifiers are used as input to an AdaBoost machine learning algorithm to train a strong classifier. In practice, we trained a cascade with 4 strong classifiers containing 79 features. An adaptive binarization and extension algorithm is applied to those regions selected by the cascade classifier. Commercial OCR software is used to read the text or reject it as a non-text region. The overall algorithm has a success rate of over 90% (evaluated by complete detection and reading of the text) on the test set and the unread text is typically small and distant from the viewer.