Why multiple document image binarizations improve OCR

Why multiple document image binarizations improve OCR
复制标题

为什么多个文档图像二值化可以改善 OCR

DOI:
--
复制
发表时间:
2013
期刊:
The Hip
影响因子:
--
通讯作者:
Eric K. Ringger
Eric K. Ringger
中科院分区:
--
文献类型:
--
作者:
William B. Lund;Douglas J. Kennard;Eric K. Ringger

文献摘要

被引文献

相似文献

我们之前的工作表明,通过使用多个信息源和多个 OCR 假设(包括来自多个文档图像二值化的假设),光学字符识别 (OCR) 对退化的历史机器打印文档的纠错得到了改善。本文的贡献在于展示了多个二值化之间的多样性如何使 OCR 准确性的提高成为可能。我们演示了校正所需的信息在给定文档图像的多个二值化中分布的程度和广度。我们的分析表明,这些更正的来源并不限于任何单个二值化,并且所有二值化都包含实现最佳结果所需的信息(通过最终 OCR 决策的字错误率 (WER) 来衡量)。即使具有高 WER 的二值化也有助于改善最终的 OCR。对于本研究中使用的语料库,使用 WER 最低的二值化图像的 OCR 中未找到的假设校正了全部标记的 2.68%。此外,我们还发现 OCR 整体的 WER 越高,在文档图像的所有二值化中分布的校正就越多。
Our previous work has shown that the error correction of optical character recognition (OCR) on degraded historical machine-printed documents is improved with the use of multiple information sources and multiple OCR hypotheses including from multiple document image binarizations. The contributions of this paper are in demonstrating how diversity among multiple binarizations makes those improvements to OCR accuracy possible. We demonstrate the degree and breadth to which the information required for correction is distributed across multiple binarizations of a given document image. Our analysis reveals that the sources of these corrections are not limited to any single binarization and that the full range of binarizations holds information needed to achieve the best result as measured by the word error rate (WER) of the final OCR decision. Even binarizations with high WERs contribute to improving the final OCR. For the corpus used in this research, fully 2.68% of all tokens are corrected using hypotheses not found in the OCR of the binarized image with the lowest WER. Further, we show that the higher the WER of the OCR overall, the more the corrections are distributed among all binarizations of the document image.