Multi-modular domain-tailored OCR post-correction

Multi-modular domain-tailored OCR post-correction
复制标题

多模块领域定制 OCR 后校正

DOI:
10.18653/v1/d17-1288
复制
发表时间:
2017
期刊:
ArXiv
影响因子:
--
通讯作者:
Jonas Kuhn
Jonas Kuhn
中科院分区:
--
文献类型:
--
作者:
Sarah Schulz;Jonas Kuhn

文献摘要

被引文献

相似文献

许多数字人文项目的主要障碍之一是低数据可用性。文本数字化是一个昂贵且耗时的过程,而光学字符识别(OCR)后校正是时间关键因素之一。在OCR后校正的例子中,我们展示了一个通用系统的适应性,以解决一个特定的问题,只有很少的数据。该系统解释了在文学领域中来自不同时期的OCRed文本中遇到的各种错误。我们展示了不同方法的组合,例如统计机器翻译和拼写检查,在排名机制的帮助下,比单一方法有了巨大的提高。由于我们认为结果工具的可访问性是数字人文学科合作的关键部分,因此我们描述了我们建议的高效文本识别和随后的自动和手动后期校正的工作流程
One of the main obstacles for many Digital Humanities projects is the low data availability. Texts have to be digitized in an expensive and time consuming process whereas Optical Character Recognition (OCR) post-correction is one of the time-critical factors. At the example of OCR post-correction, we show the adaptation of a generic system to solve a specific problem with little data. The system accounts for a diversity of errors encountered in OCRed texts coming from different time periods in the domain of literature. We show that the combination of different approaches, such as e.g. Statistical Machine Translation and spell checking, with the help of a ranking mechanism tremendously improves over single-handed approaches. Since we consider the accessibility of the resulting tool as a crucial part of Digital Humanities collaborations, we describe the workflow we suggest for efficient text recognition and subsequent automatic and manual post-correction