Font and Function Word Identification in Document Recognition

Font and Function Word Identification in Document Recognition
复制标题

文档识别中的字体和功能词识别

DOI:
10.1006/cviu.1996.0005
复制
发表时间:
1996
期刊:
Comput. Vis. Image Underst.
影响因子:
--
通讯作者:
J. Hull
J. Hull
中科院分区:
--
文献类型:
--
作者:
S. Khoubyari;J. Hull

文献摘要

被引文献

相似文献

提出了一种算法,确定在英语文档中的运行文本打印的主要字体。常见的功能词(如the、of、and、a和to)也被识别为字体标识的一部分。从输入文档生成词图像的集群,并将其与从字体和文档图像导出的功能词的数据库进行匹配。最匹配的字体或文档提供了主要字体和功能词的标识。这种技术利用了这样一个事实,即大多数机器打印的文档都是用一种主要字体准备的。此外,利用文档中的重复单词来克服输入中的噪声。该技术的优点包括其用作文档识别算法的预处理步骤。实验结果表明,高精度的原始和退化的文档图像数据库上实现。
An algorithm is presented that identifies the predominant font in which the running text in an English language document is printed. Frequent function words (such asthe,of,and,a, andto) are also recognized as part of the font identification. Clusters of word images are generated from an input document and matched to a database of function words derived from fonts and document images. The font or document that matches best provides the identification of the predominant font and function words. This technique takes advantage of the fact that most machine-printed documents are prepared with a single predominant font. Also, the repeated words in the document are utilized to overcome noise in the input. Advantages of this technique include its use as a preprocessing step for a document recognition algorithm. Experimental results show high accuracy is achieved on a database of original and degraded document images.