Crowdsourcing an OCR Gold Standard for a German and French Heritage Corpus

Crowdsourcing an OCR Gold Standard for a German and French Heritage Corpus
复制标题

众包德国和法国遗产语料库的 OCR 黄金标准

DOI:
--
复制
发表时间:
2016
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
M. Volk
M. Volk
中科院分区:
--
文献类型:
--
作者:
S. Clematide;Lenz Furrer;M. Volk

文献摘要

被引文献

相似文献

OCR输出(光学字符识别)后校正的众包方法已成功应用于几个历史文本集。我们报告了我们的人群校正平台Kokos,我们建立该平台是为了提高19世纪瑞士阿尔卑斯山俱乐部(SAC)数字化年鉴的OCR质量。这一多语种遗产语料库由主要用德语和法语书写的阿尔卑斯语文本组成,所有文本均采用Antiqua字体。寻找并吸引志愿者将大量页面更正为高质量的文本需要精心设计的用户界面,易于使用的工作流程以及保持参与者积极性的持续努力。在大约7个月的时间里,志愿者在大约21,000页中纠正了超过180,000个字符,达到了OCR黄金标准,在单词水平上系统评估的准确率为99.7%。众包OCR黄金标准和来自Abby FineReader 7的每页相应原始OCR识别结果可作为资源提供。此外,所有页面的扫描图像(300 dpi)都包括在内,以方便使用其他OCR软件进行测试。
Crowdsourcing approaches for post-correction of OCR output (Optical Character Recognition) have been successfully applied to several historic text collections. We report on our crowd-correction platform Kokos, which we built to improve the OCR quality of the digitized yearbooks of the Swiss Alpine Club (SAC) from the 19th century. This multilingual heritage corpus consists of Alpine texts mainly written in German and French, all typeset in Antiqua font. Finding and engaging volunteers for correcting large amounts of pages into high quality text requires a carefully designed user interface, an easy-to-use workflow, and continuous efforts for keeping the participants motivated. More than 180,000 characters on about 21,000 pages were corrected by volunteers in about 7 month, achieving an OCR gold standard with a systematically evaluated accuracy of 99.7% on the word level. The crowdsourced OCR gold standard and the corresponding original OCR recognition results from Abby FineReader 7 for each page are available as a resource. Additionally, the scanned images (300dpi) of all pages are included in order to facilitate tests with other OCR software.