A Fast Scene Text Detector Using Knowledge Distillation

A Fast Scene Text Detector Using Knowledge Distillation
复制标题

使用知识蒸馏的快速场景文本检测器

DOI:
10.1109/access.2019.2895330
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Yang, Guowei
Yang, Guowei
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yang, Peng;Zhang, Fanlong;Yang, Guowei

文献摘要

被引文献

相似文献

由于自然图像中文本的任意方向、低分辨率、透视失真和不同的长宽比,附带场景文本检测是一个具有挑战性的问题。在本文中,我们提出了一种端到端的可训练深度模型,可以有效且高效地定位多方向场景文本。我们的检测器包括学生网络和教师网络,它们分别继承了复杂的 VGGNet 和轻量级的 PVANet 架构。在部署文本检测时,教师网络用于通过知识提炼来指导学生的训练过程,以保持准确性和效率之间的权衡。我们在三个流行的基准上评估了所提出的检测器,它在 ICDAR2015 Incidental Scene Text、COCO-Text 和 ICDAR2013 上分别实现了 83.7%、57.27% 和 90% 的 F 测量,优于最先进的方法。
Incidental scene text detection is a challenging problem because of arbitrary orientation, low resolution, perspective distortion, and variant aspect ratios of text in natural images. In this paper, we present an end-to-end trainable deep model, which can effectively and efficiently locate multi-oriented scene text. Our detector includes a student network and a teacher network, and they inherit complex VGGNet and lightweight PVANet architecture, respectively. While deploying for text detection, the teacher network is used to guide the training process of a student via knowledge distilling so as to maintain the tradeoff between accuracy and efficiency. We have evaluated the proposed detector on three popular benchmarks, and it achieves F-measures of 83.7%, 57.27%, and 90% on ICDAR2015 Incidental Scene Text, COCO-Text, and ICDAR2013, respectively, which outperforms the most state-of-the-art methods.