OCR Post-Processing Text Correction using Simulated Annealing (OPTeCA)
OCR Post-Processing Text Correction using Simulated Annealing (OPTeCA)
复制标题
使用模拟退火 (OPTeCA) 进行 OCR 后处理文本校正
DOI:
--
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
G. Khirbat
中科院分区:
文献类型:
--
作者:
G. Khirbat
This paper describes the system details and results of team “EOF” from the University of Melbourne for the shared task of ALTA 2017, which addresses the problem of text correction for post-processed Optical Character Recognition (OCR) based systems. We developed a two stage sys-tem which first detects errors in the given OCR post-processed text with the help of a support vector machine trained using given training dataset, followed by rectifying the errors by employing a confidence-based mechanism using simulated annealing to obtain an optimal correction from a pool of candidate corrections. Our system achieved a F 1 -score of 32.98% on the private leaderboard 1 , which is the best score among all the participating systems.