LAL: Linguistically Aware Learning for Scene Text Recognition

LAL: Linguistically Aware Learning for Scene Text Recognition
复制标题

DOI:
10.1145/3394171.3413913
复制
发表时间:
2020-10
期刊:
Proceedings of the 28th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Y. Zheng;Wenda Qin;D. Wijaya;Margrit Betke
Y. Zheng;Wenda Qin;D. Wijaya;Margrit Betke
中科院分区:
其他
文献类型:
--
作者:
Y. Zheng;Wenda Qin;D. Wijaya;Margrit Betke

文献摘要

被引文献

相似文献

场景文本识别是对自然场景图像中的字符序列进行识别的任务。在场景图像和潜在的高度复杂的背景中文本外观的相当大的多样性使文本识别具有挑战性。以前的方法使用字符序列生成器来分析文本区域,然后将候选字符序列与语言模型进行比较。在这项工作中,我们提出了一个双峰框架,同时利用视觉和语言信息来提高识别性能。我们的语言感知学习(LAL)方法使用整流、编码器和注意解码器方法有效地学习视觉嵌入,使用深度下一个字符预测模型有效地学习语言嵌入。我们提出了一种有效结合这两种嵌入的创新方法。我们在八个标准基准上的实验表明,我们的方法在很大程度上优于以前的方法,特别是在旋转、缩短和弯曲的文本上。我们表明,双峰方法具有统计上显著的影响。我们还提供了一个新的数据集,并在流水线文本识别框架中将LAL与文本检测器结合使用时显示出强大的性能。
Scene text recognition is the task of recognizing character sequences in images of natural scenes. The considerable diversity in the appearance of text in a scene image and potentially highly complex backgrounds make text recognition challenging. Previous approaches employ character sequence generators to analyze text regions and, subsequently, compare the candidate character sequences against a language model. In this work, we propose a bimodal framework that simultaneously utilizes visual and linguistic information to enhance recognition performance. Our linguistically aware learning (LAL) method effectively learns visual embeddings using a rectifier, encoder, and attention decoder approach, and linguistic embeddings, using a deep next-character prediction model. We present an innovative way of combining these two embeddings effectively. Our experiments on eight standard benchmarks show that our method outperforms previous methods by large margins, particularly on rotated, foreshortened, and curved text. We show that the bimodal approach has a statistically significant impact. We also contribute a new dataset, and show robust performance when LAL is combined with a text detector in a pipelined text spotting framework.