Font and Function Word Identification in Document Recognition
Font and Function Word Identification in Document Recognition
复制标题
文档识别中的字体和功能词识别
DOI:
10.1006/cviu.1996.0005
复制
发表时间:
1996
期刊:
影响因子:
--
通讯作者:
J. Hull
中科院分区:
文献类型:
--
作者:
S. Khoubyari;J. Hull
An algorithm is presented that identifies the predominant font in which the running text in an English language document is printed. Frequent function words (such asthe,of,and,a, andto) are also recognized as part of the font identification. Clusters of word images are generated from an input document and matched to a database of function words derived from fonts and document images. The font or document that matches best provides the identification of the predominant font and function words. This technique takes advantage of the fact that most machine-printed documents are prepared with a single predominant font. Also, the repeated words in the document are utilized to overcome noise in the input. Advantages of this technique include its use as a preprocessing step for a document recognition algorithm. Experimental results show high accuracy is achieved on a database of original and degraded document images.