Ensemble Optical Character Recognition Systems via Machine Learning

Ensemble Optical Character Recognition Systems via Machine Learning
复制标题

通过机器学习集成光学字符识别系统

DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Haowen Cao
Haowen Cao
中科院分区:
--
文献类型:
--
作者:
Zifei Shan;Haowen Cao

文献摘要

被引文献

相似文献

光学字符识别 (OCR) 系统广泛用于将扫描文本处理为计算机可用的文本。我们观察到当前的 OCR 系统在特定领域的论文上表现不佳,甚至会生成大量错误的单词;此外,不同的OCR系统会出现相对独立的错误。基于这些观察,我们从多个开源 OCR 系统中训练一个集成系统,该系统从每个 OCR 生成的候选中选择输出,并使用机器学习技术训练该系统。我们实现了 Softmax 回归和多类 SVM。我们的系统在 1,011 个单词的训练集上的不同输出之间进行选择的准确率超过 80%。我们通过提出新选项来进一步探索提高性能的方法,并利用领域知识来提高其性能。我们的贡献在于以下几个方面:(1)我们展示了将 OCR 系统视为黑匣子并相互纠正其输出的巨大潜力。 (2) 我们的系统建立在最好的开源 OCR 的基础上,并在准确性上取得了显着的提高。 (3) 此外,我们的工作探索了利用丰富的语义知识来构建更好的 OCR 系统的可能性,并深入了解将集成系统作为黑匣子的通用方法。
Optical Character Recognition (OCR) Systems are widely used to process scanned text into text usable by computers. We observe that current OCR systems have bad performance on domain-specific papers, even generating lots of incorrect words; besides, different OCR systems make relatively independent mistakes. Based on these observations, we train an ensemble system from multiple open-source OCR systems, which chooses outputs from candidates generated by each OCR, and train the system with machine learning techniques. We implement Softmax Regression and multi-class SVM. Our system achieve over 80% accuracy selecting between different outputs on our training set of 1,011 words. We further explore ways to improve the performance by suggesting new options, and use domain knowledge to improve its performance. Our contribution lies in following aspects: (1) We show the great potential of treating OCR systems as black-boxes and correct their outputs from each other. (2) Our system build on best open-source OCRs and achieve significant improvement on their accuracy. (3) Moreover, our work explore the possibility to make use of rich semantic knowledge to craft a better OCR system, and cast insight to a general approach to ensemble systems as black-boxes.