Parallel Scale-wise Attention Network for Effective Scene Text Recognition

Parallel Scale-wise Attention Network for Effective Scene Text Recognition
复制标题

DOI:
10.1109/ijcnn52387.2021.9534223
复制
发表时间:
2021-04
期刊:
2021 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Usman Sajid;Michael Chow;Jin Zhang;Taejoon Kim;Guanghui Wang
Usman Sajid;Michael Chow;Jin Zhang;Taejoon Kim;Guanghui Wang
中科院分区:
其他
文献类型:
--
作者:
Usman Sajid;Michael Chow;Jin Zhang;Taejoon Kim;Guanghui Wang

文献摘要

相似文献

本文提出了一种新的场景文本图像的文本识别网络。许多最先进的方法在文本编码器或解码器中使用注意机制来进行文本对齐。虽然基于编码器的注意力产生有希望的结果,这些计划继承了显着的局限性。它们依次执行特征提取(FE)和视觉注意(VA),这将注意力机制限制为仅依赖于FE最终单尺度输出。此外,注意过程的利用仅限于直接将其应用于单尺度特征图。为了解决这些问题,我们提出了一种新的多尺度和基于编码器的注意力网络,用于并行执行多尺度FE和VA的文本识别。多尺度通道之间也会进行规律性的融合,共同发展协调的知识。标准基准测试的定量评估和鲁棒性分析表明,在大多数情况下,该网络的性能优于最先进的。
The paper proposes a new text recognition network for scene-text images. Many state-of-the-art methods employ the attention mechanism either in the text encoder or decoder for the text alignment. Although the encoder-based attention yields promising results, these schemes inherit noticeable limitations. They perform the feature extraction (FE) and visual attention (VA) sequentially, which bounds the attention mechanism to rely only on the FE final single-scale output. Moreover, the utilization of the attention process is limited by only applying it directly to the single scale feature-maps. To address these issues, we propose a new multi-scale and encoder-based attention network for text recognition that performs the multi-scale FE and VA in parallel. The multi-scale channels also undergo regular fusion with each other to develop the coordinated knowledge together. Quantitative evaluation and robustness analysis on the standard benchmarks demonstrate that the proposed network outperforms the state-of-the-art in most cases.