Indexing and retrieval of words in old documents

Indexing and retrieval of words in old documents
复制标题

旧文档中单词的索引和检索

DOI:
10.1109/icdar.2003.1227663
复制
发表时间:
2003
期刊:
Seventh International Conference on Document Analysis and Recognition, 2003. Proceedings.
影响因子:
--
通讯作者:
G. Soda
G. Soda
中科院分区:
--
文献类型:
--
作者:
S. Marinai;E. Marino;G. Soda

文献摘要

被引文献

相似文献

本文介绍了一个系统,有效的索引和检索的话,在收集的文件图像。所提出的方法是基于两个主要原则:无监督的原型聚类,和字符串编码的高效字符串匹配。在索引期间,训练自组织映射(SOM),以便将待存储的文档的子集中的类似符号(类似字符的对象)聚类在一起。通过使用经过训练的SOM,整个集合中的单词可以被存储并用固定长度的描述表示,该描述可以很容易地进行比较,以便响应于用户查询对最相似的单词进行评分。该系统可以自动适应不同的语言和字体风格。最合适的应用程序是处理旧文件(18世纪和19世纪),而当前的OCR面临更多困难。实验结果描述了三个应用场景具有不同的难度水平,目前的OCR系统。
This paper describes a system for efficient indexing and retrieval of words in collections of document images. The proposed method is based on two main principles: unsupervised prototype clustering, and string encoding for efficient string matching. During indexing, a self organizing map (SOM) is trained so as to cluster together similar symbols (character-like objects) in a sub-set of the documents to be stored. By using the trained SOM the words in the whole collection can be stored and represented with a fixed-length description that can be easily compared in order to score most similar words in response to a user query. The system can be automatically adapted to different languages and font styles. The most appropriate applications are for the processing of old documents (18th and 19th Centuries) where current OCRs have more difficulties. Experimental results describe three application scenarios having various levels of difficulty for current OCR systems.