ICDAR2017 Competition on Post-OCR Text Correction

ICDAR2017 Competition on Post-OCR Text Correction
复制标题

ICDAR2017 OCR后文本校正竞赛

DOI:
10.1109/icdar.2017.232
复制
发表时间:
2017
期刊:
2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR)
影响因子:
--
通讯作者:
Jean
Jean
中科院分区:
--
文献类型:
--
作者:
Guillaume Chiron;A. Doucet;Mickaël Coustaty;Jean

文献摘要

被引文献

相似文献

本文介绍了ICDAR2017 OCR后文本更正竞赛,并介绍了参赛者提交的不同方法。OCR在过去的30年里一直是一个活跃的研究领域,但结果仍然不完美,特别是对于历史文献。本次比赛的目的是比较和评估纠正(去噪)OCR文本的自动方法。挑战包括两个独立的任务:1)错误检测和2)错误纠正。向参与者提供了12 M OCR符号沿着的原始数据集,其中80%的数据集用于培训,20%用于评估。汇总了不同的来源,即包含涵盖两种语言(英语和法语)的报纸和专著。11个团队提交了结果,而平均而言,只有一半的提交方法能够对评估数据集进行降噪,这突出了任务的难度。在任何情况下,这场比赛,其中计数35个注册,说明了社区对这个基本问题的强烈兴趣,这是任何涉及文本数据的数字化过程的关键。
This paper describes the ICDAR2017 competition on post-OCR text correction and presents the different methods submitted by the participants. OCR has been an active research field for over the past 30 years but results are still imperfect, especially for historical documents. The purpose of this competition is to compare and evaluate automatic approaches for correcting (denoising) OCR-ed texts. The challenge consists of two independent tasks: 1) error detection and 2) error correction. An original dataset of 12M OCR-ed symbols along with an aligned ground truth was provided to the participants with 80% of the dataset dedicated to the training and 20% to the evaluation. Different sources were aggregated and namely contain newspapers and monographs covering 2 languages (English and French). 11 teams submitted results, while the difficulty of the task was underlined by the fact that only half of the submitted methods were able to denoise the evaluation dataset on average. In any case, this competition, which counted 35 registrations, illustrates the strong interest of the community in this essential problem, which is key to any digitization process involving textual data.