Crowdsourcing in the Document Processing Practice - (A Short Practitioner/Visionary Paper)

Crowdsourcing in the Document Processing Practice - (A Short Practitioner/Visionary Paper)
复制标题

文档处理实践中的众包 - (简短的实践者/有远见的论文)

DOI:
--
复制
发表时间:
2010
期刊:
ICWE Workshops
影响因子:
--
通讯作者:
Tal Drory
Tal Drory
中科院分区:
--
文献类型:
--
作者:
E. Karnin;E. Walach;Tal Drory

文献摘要

被引文献

相似文献

扫描文档的处理要求通过OCR(光学字符识别)计算机程序自动识别文本,然后进行人工验证和校正。这些基本手动任务的众包是一个很好的选择,前提是可以处理一些关键挑战,以便满足客户期望的质量水平。我们展示了如何调整和增强有效验证和纠正工具,以解决与众包相关的问题,如数据隐私,质量控制,人群监控和工作质量保证。我们开始在我们的COOPERATION ENGINE for Correction of Extracted Text(CONCERT)中实现这些想法和技术,该引擎用于图书数字化项目。
The processing of scanned documents calls for automatic recognition of the text by OCR (Optical Character Recognition) computer programs, followed by human validation and correction. Crowdsourcing of these essential manual tasks is a good option, provided one can take care of some key challenges, so that the quality level expected by the customer is met. We show how tools for efficient validation and correction are adapted and enhanced to address issues associated with crowdsourcing, such as data privacy, quality control, crowd monitoring, and job quality assurance. We started to implement these ideas and technologies in our COoperative eNgine for Correction of ExtRacted Text (CONCERT), which is used in book digitization projects.