Instance Segmentation Network With Self-Distillation for Scene Text Detection

Instance Segmentation Network With Self-Distillation for Scene Text Detection
复制标题

基于自蒸馏的场景文本检测实例分割网络

DOI:
10.1109/access.2020.2978225
复制
发表时间:
2020-03
期刊:
影响因子:
3.9
通讯作者:
Peng Yang;Guowei Yang;Xun Gong;Pingping Wu;Xu Han;Jiasong Wu;Caisen Chen
Peng Yang;Guowei Yang;Xun Gong;Pingping Wu;Xu Han;Jiasong Wu;Caisen Chen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Peng Yang;Guowei Yang;Xun Gong;Pingping Wu;Xu Han;Jiasong Wu;Caisen Chen

文献摘要

相似文献

基于图像分割的文本检测方法已经成为检测任意方向和形状的场景文本的主流方法。然而,为了解决诸如分离彼此非常接近的文本实例等具有挑战性的问题,这些方法通常需要耗时的后处理。在本文中,我们提出了一个实例分割网络(ISNet),它同时生成原型掩码和每个实例掩码系数。在线性组合这两个组件后,ISNet可以实现快速文本定位。此外,我们应用自蒸馏来训练ISNet并提高其检测精度。我们已经在四个流行的基准上评估了所提出的方法,即,ICDAR 2015、ICDAR 2017 MLT、CTW 1500和Total-Text的实验结果表明,该算法能够在场景文本检测的准确性和效率之间取得较好的折衷。
Segmentation based methods have become the mainstream for detecting scene text with arbitrary orientations and shapes. In order to address challenging problems such as separating the text instances that are very close to each other, however, these methods often require time-consuming post-processing. In this paper, we propose an instance segmentation network (ISNet), which simultaneously generates prototype masks and per-instance mask coefficients. After linearly combining the two components, ISNet can implement fast text location. Furthermore, we apply self-distillation to train the ISNet and refine its detection accuracy. We have evaluated the proposed method on four popular benchmarks, i.e., ICDAR2015, ICDAR2017 MLT, CTW1500 and Total-Text, and the experimental results show that it can achieve better tradeoff between accuracy and efficiency for scene text detection.