Ensemble Optical Character Recognition Systems via Machine Learning
Ensemble Optical Character Recognition Systems via Machine Learning
复制标题
通过机器学习集成光学字符识别系统
DOI:
--
复制
发表时间:
2013
期刊:
影响因子:
--
通讯作者:
Haowen Cao
中科院分区:
文献类型:
--
作者:
Zifei Shan;Haowen Cao
Optical Character Recognition (OCR) Systems are widely used to process scanned text into text usable by computers. We observe that current OCR systems have bad performance on domain-specific papers, even generating lots of incorrect words; besides, different OCR systems make relatively independent mistakes. Based on these observations, we train an ensemble system from multiple open-source OCR systems, which chooses outputs from candidates generated by each OCR, and train the system with machine learning techniques. We implement Softmax Regression and multi-class SVM. Our system achieve over 80% accuracy selecting between different outputs on our training set of 1,011 words. We further explore ways to improve the performance by suggesting new options, and use domain knowledge to improve its performance. Our contribution lies in following aspects: (1) We show the great potential of treating OCR systems as black-boxes and correct their outputs from each other. (2) Our system build on best open-source OCRs and achieve significant improvement on their accuracy. (3) Moreover, our work explore the possibility to make use of rich semantic knowledge to craft a better OCR system, and cast insight to a general approach to ensemble systems as black-boxes.