Cleaning Dirty Books: Post-OCR Processing for Previously Scanned Texts

Cleaning Dirty Books: Post-OCR Processing for Previously Scanned Texts
复制标题

DOI:
10.18653/v1/2021.findings-emnlp.356
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Allen Kim;Charuta G. Pethe;Naoya Inoue;S. Skiena
Allen Kim;Charuta G. Pethe;Naoya Inoue;S. Skiena
中科院分区:
其他
文献类型:
--
作者:
Allen Kim;Charuta G. Pethe;Naoya Inoue;S. Skiena

文献摘要

相似文献

为了进行NLP分析,需要清理大量的数字化书籍,这是因为扫描文本中存在错误,语料库中存在重复的卷。在本文中,我们考虑在存在光学字符识别(OCR)错误的情况下重复数据删除的问题。我们提出了处理这些错误的方法,对来自古滕贝格项目数据集的19,347个文本和来自HathiTrust图书馆的96,635个文本进行了评估。我们证明了语言模型的改进现在可以在不考虑扫描图像本身的情况下检测和纠正OCR错误。通过对齐相同基础工作的扫描对发现的不一致性提供了训练数据,以构建用于检测和纠正错误的模型。我们从58,808次扫描中识别出17,136本重复扫描的书籍中的每一本的规范版本。最后,我们研究的方法来检测和纠正单拷贝文本中的错误。我们表明,平均而言,我们的方法纠正的错误是它引入的错误的六倍以上。我们还提供了有趣的分析扫描质量和其他因素,如位置和出版年份之间的关系。
Substantial amounts of work are required to clean large collections of digitized books for NLP analysis, both because of the presence of errors in the scanned text and the presence of duplicate volumes in the corpora. In this paper, we consider the issue of deduplication in the presence of optical character recognition (OCR) errors. We present methods to handle these errors, evaluated on a collection of 19,347 texts from the Project Gutenberg dataset and 96,635 texts from the HathiTrust Library. We demonstrate that improvements in language models now enable the detection and correction of OCR errors without consideration of the scanning image itself. The inconsistencies found by aligning pairs of scans of the same underlying work provides training data to build models for detecting and correcting errors. We identify the canonical version for each of 17,136 repeatedly-scanned books from 58,808 scans. Finally, we investigate methods to detect and correct errors in single-copy texts. We show that on average, our method corrects over six times as many errors as it introduces. We also provide interesting analysis on the relation between scanning quality and other factors such as location and publication year.