Neural OCR Post-Hoc Correction of Historical Corpora

Neural OCR Post-Hoc Correction of Historical Corpora
复制标题

DOI:
10.1162/tacl_a_00379
复制
发表时间:
2021-02
影响因子:
10.9
通讯作者:
Lijun Lyu;Maria Koutraki;Martin Krickl;B. Fetahu
Lijun Lyu;Maria Koutraki;Martin Krickl;B. Fetahu
中科院分区:
人文科学1区
文献类型:
--
作者:
Lijun Lyu;Maria Koutraki;Martin Krickl;B. Fetahu

文献摘要

相似文献

摘要 光学字符识别 (OCR) 对于更深入地访问历史馆藏至关重要。 OCR 需要考虑正字法变化、字体或语言演变(即新字母、单词拼写),它们是字符、单词或分词转录错误的主要来源。对于历史印刷品的数字语料库,由于扫描质量低和缺乏语言标准化,错误进一步加剧。对于 OCR 事后纠正任务,我们提出了一种基于循环网络 (RNN) 和深度卷积网络 (ConvNet) 相结合的神经方法来纠正 OCR 转录错误。在字符级别,我们灵活地捕获错误,并基于新颖的注意力机制解码正确的输出。考虑到输入和输出的相似性,我们提出了一个新的损失函数来奖励模型的纠正行为。对德语历史书籍语料库的评估表明,我们的模型在捕获各种 OCR 转录错误方面具有鲁棒性,并将 32.3% 的单词错误率降低了 89% 以上。
Abstract Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of character, word, or word segmentation transcription errors. For digital corpora of historical prints, the errors are further exacerbated due to low scan quality and lack of language standardization. For the task of OCR post-hoc correction, we propose a neural approach based on a combination of recurrent (RNN) and deep convolutional network (ConvNet) to correct OCR transcription errors. At character level we flexibly capture errors, and decode the corrected output based on a novel attention mechanism. Accounting for the input and output similarity, we propose a new loss function that rewards the model’s correcting behavior. Evaluation on a historical book corpus in German language shows that our models are robust in capturing diverse OCR transcription errors and reduce the word error rate of 32.3% by more than 89%.