Approximate matching for OCR-processed bibliographic data

Approximate matching for OCR-processed bibliographic data
复制标题

OCR 处理的书目数据的近似匹配

DOI:
10.1109/icpr.1996.546933
复制
发表时间:
1996
期刊:
Proceedings of 13th International Conference on Pattern Recognition
影响因子:
--
通讯作者:
J. Adachi
J. Adachi
中科院分区:
--
文献类型:
--
作者:
A. Takasu;Norio Katayama;M. Yamaoka;O. Iwaki;K. Oyama;J. Adachi

文献摘要

被引文献

相似文献

本文提出了一种将以文档图像形式获取的学术论文参考文献中的书目与书目数据库记录进行匹配的方法。本文的主要主题是对文献理解方法论所获得的错误书目数据进行处理。该方法通过近似匹配从参考文献和数据库中的书目数据串中选择k个子串进行精确匹配,从而在不考虑字符串错误的情况下从参考数据库中找到候选记录集。对于OCR的精度/SPLα/,理论观察表明,在假设OCR误差在字符串中随机且独立地发生的情况下,该方法的精度为1-(1-/SPLα//sup m/)/sup k/。将该方法应用于187篇日文参考文献,取得了94.05%的准确率。
This paper presents a method for matching bibliographies in references of academic papers obtained as document images with records of bibliographic databases. The main subject of this paper is to handle the erroneous bibliographic data obtained by a document understanding methodology. The presented method can find a candidate record set from referral databases in spite of the errors of string by means of approximate matching which is performed as an exact matching of k substrings of length m chosen from the strings of bibliographic data in references and in databases. For the accuracy /spl alpha/ of the OCR, theoretical observation shows that the accuracy of the presented method is 1-(1-/spl alpha//sup m/)/sup k/ under the assumption that the OCR error occurs randomly and independently in the string. The method is applied to references of 187 Japanese articles and achieves accuracy of 94.05%.