BART for Post-Correction of OCR Newspaper Text

BART for Post-Correction of OCR Newspaper Text
复制标题

BART 用于 OCR 报纸文本的后期校正

DOI:
--
复制
发表时间:
2021
期刊:
WNUT
影响因子:
--
通讯作者:
Yen
Yen
中科院分区:
--
文献类型:
--
作者:
Elizabeth Soper;Stanley Fujimoto;Yen

文献摘要

被引文献

相似文献

由于旧文档的退化和排版的变化,报纸页面图像的光学字符识别(OCR)容易受到噪声的影响。在这份报告中,我们提出了一种新的OCR后校正方法。我们将纠错作为一项翻译任务,并对BART进行了微调,BART是一种基于转换器的序列到序列语言模型,经过预先训练,可以对损坏的文本进行去噪。我们首次将句子级转换模型用于OCR后校正,并且我们最好的模型在字符准确率上比原始的噪声OCR文本提高了29.4%。我们的结果证明了预先训练的语言模型在处理噪声文本方面的有效性。
Optical character recognition (OCR) from newspaper page images is susceptible to noise due to degradation of old documents and variation in typesetting. In this report, we present a novel approach to OCR post-correction. We cast error correction as a translation task, and fine-tune BART, a transformer-based sequence-to-sequence language model pretrained to denoise corrupted text. We are the first to use sentence-level transformer models for OCR post-correction, and our best model achieves a 29.4% improvement in character accuracy over the original noisy OCR text. Our results demonstrate the utility of pretrained language models for dealing with noisy text.