OCR Post-Processing Text Correction using Simulated Annealing (OPTeCA)

OCR Post-Processing Text Correction using Simulated Annealing (OPTeCA)
复制标题

使用模拟退火 (OPTeCA) 进行 OCR 后处理文本校正

DOI:
--
复制
发表时间:
2017
期刊:
Australasian Language Technology Association Workshop
影响因子:
--
通讯作者:
G. Khirbat
G. Khirbat
中科院分区:
--
文献类型:
--
作者:
G. Khirbat

文献摘要

被引文献

相似文献

本文介绍了来自墨尔本大学的团队“ZH”在阿尔塔2017年的共享任务中的系统细节和结果,该任务解决了基于后处理光学字符识别(OCR)系统的文本校正问题。我们开发了一个两阶段系统,首先在使用给定训练数据集训练的支持向量机的帮助下检测给定OCR后处理文本中的错误,然后通过采用基于置信度的机制使用模拟退火来纠正错误,以从候选校正池中获得最佳校正。我们的系统在私人排行榜1上获得了32.98%的F1分数,这是所有参与系统中最好的分数。
This paper describes the system details and results of team “EOF” from the University of Melbourne for the shared task of ALTA 2017, which addresses the problem of text correction for post-processed Optical Character Recognition (OCR) based systems. We developed a two stage sys-tem which first detects errors in the given OCR post-processed text with the help of a support vector machine trained using given training dataset, followed by rectifying the errors by employing a confidence-based mechanism using simulated annealing to obtain an optimal correction from a pool of candidate corrections. Our system achieved a F 1 -score of 32.98% on the private leaderboard 1 , which is the best score among all the participating systems.