PP-OCR: A Practical Ultra Lightweight OCR System

PP-OCR: A Practical Ultra Lightweight OCR System
复制标题

DOI:
--
复制
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuning Du;Chenxia Li;Ruoyu Guo;Xiaoting Yin;Weiwei Liu;Jun Zhou;Yifan Bai;Zilin Yu;Yehua Yang;Qingqing Dang;Hongya Wang
Yuning Du;Chenxia Li;Ruoyu Guo;Xiaoting Yin;Weiwei Liu;Jun Zhou;Yifan Bai;Zilin Yu;Yehua Yang;Qingqing Dang;Hongya Wang
中科院分区:
其他
文献类型:
--
作者:
Yuning Du;Chenxia Li;Ruoyu Guo;Xiaoting Yin;Weiwei Liu;Jun Zhou;Yifan Bai;Zilin Yu;Yehua Yang;Qingqing Dang;Hongya Wang

文献摘要

被引文献

相似文献

光学字符识别(OCR)系统已广泛应用于办公自动化(OA)系统、工厂自动化、在线教育、地图制作等各种应用场景中。然而,由于文本外观的多样性和对计算效率的要求,OCR仍然是一项具有挑战性的任务。在本文中,我们提出了一个实用的超轻量级OCR系统,即,PP-OCR。PP-OCR的整体模型大小仅为3.5M,用于识别6622个汉字和2.8M,用于识别63个字母数字符号。我们引入了一系列策略来增强模型能力或减小模型大小。文中还提供了相应的烧蚀实验和真实的数据。同时,发布了几个用于中文和英文识别的预训练模型,包括文本检测器(使用97 K图像),方向分类器(使用600 K图像)以及文本识别器(使用17.9M图像)。此外,本文还在法语、韩语、日语和德语等多种语言识别任务中对PP-OCR进行了验证。上述所有模型都是开源的,代码可以在GitHub存储库中找到,即,https URL。
The Optical Character Recognition (OCR) systems have been widely used in various of application scenarios, such as office automation (OA) systems, factory automations, online educations, map productions etc. However, OCR is still a challenging task due to the various of text appearances and the demand of computational efficiency. In this paper, we propose a practical ultra lightweight OCR system, i.e., PP-OCR. The overall model size of the PP-OCR is only 3.5M for recognizing 6622 Chinese characters and 2.8M for recognizing 63 alphanumeric symbols, respectively. We introduce a bag of strategies to either enhance the model ability or reduce the model size. The corresponding ablation experiments with the real data are also provided. Meanwhile, several pre-trained models for the Chinese and English recognition are released, including a text detector (97K images are used), a direction classifier (600K images are used) as well as a text recognizer (17.9M images are used). Besides, the proposed PP-OCR are also verified in several other language recognition tasks, including French, Korean, Japanese and German. All of the above mentioned models are open-sourced and the codes are available in the GitHub repository, i.e., this https URL.