ICDAR 2019 Competition on Post-OCR Text Correction

ICDAR 2019 Competition on Post-OCR Text Correction
复制标题

ICDAR 2019 OCR 后文本校正竞赛

DOI:
10.1109/icdar.2019.00255
复制
发表时间:
2017
期刊:
2019 International Conference on Document Analysis and Recognition (ICDAR)
影响因子:
--
通讯作者:
Jean
Jean
中科院分区:
--
文献类型:
--
作者:
Christophe Rigaud;A. Doucet;Mickaël Coustaty;Jean

文献摘要

被引文献

相似文献

本文介绍了ICDAR 2019年第二轮OCR后文本校正竞赛,并介绍了参赛者提交的不同方法。OCR在过去的30年里一直是一个活跃的研究领域,但结果仍然不完美,特别是对于历史文献。本次比赛的目的是比较和评估纠正(去噪)OCR文本的自动方法。当前的挑战包括两个任务:1)错误检测和2)错误校正。向参与者提供了22 M OCR符号的原始数据集(沿着)以及对齐的地面实况,其中80%的数据集用于训练,20%用于评估。对不同来源进行了汇总,其中包括报纸、历史印刷文件以及手稿和购物收据,涵盖10种欧洲语言(保加利亚语、捷克语、荷兰语、英语、芬兰语、法语、德语、波兰语、西班牙语和斯洛伐克语)。五个团队提交了结果,错误检测分数从41%到95%不等,最好的纠错改进是44%。本次比赛共有34个注册,表明了社区对改善OCR输出的强烈兴趣,这是涉及文本数据的任何数字化过程的关键问题。
This paper describes the second round of the ICDAR 2019 competition on post-OCR text correction and presents the different methods submitted by the participants. OCR has been an active research field for over the past 30 years but results are still imperfect, especially for historical documents. The purpose of this competition is to compare and evaluate automatic approaches for correcting (denoising) OCR-ed texts. The present challenge consists of two tasks: 1) error detection and 2) error correction. An original dataset of 22M OCR-ed symbols along with an aligned ground truth was provided to the participants with 80% of the dataset dedicated to training and 20% to evaluation. Different sources were aggregated and contain newspapers, historical printed documents as well as manuscripts and shopping receipts, covering 10 European languages (Bulgarian, Czech, Dutch, English, Finish, French, German, Polish, Spanish and Slovak). Five teams submitted results, the error detection scores vary from 41 to 95% and the best error correction improvement is 44%. This competition, which counted 34 registrations, illustrates the strong interest of the community to improve OCR output, which is a key issue to any digitization process involving textual data.