Recognizing Chinese Texts with Multi-width Feature Extractor and Attention-based Fusion

Recognizing Chinese Texts with Multi-width Feature Extractor and Attention-based Fusion
复制标题

使用多宽度特征提取器和基于注意力的融合识别中文文本

DOI:
10.1109/cisai54367.2021.00067
复制
发表时间:
2021
期刊:
2021 International Conference on Computer Information Science and Artificial Intelligence (CISAI)
影响因子:
--
通讯作者:
Pu Cao
Pu Cao
中科院分区:
--
文献类型:
--
作者:
Pu Cao

文献摘要

被引文献

相似文献

识别文本在光学字符识别 (OCR) 中起着重要作用。以往的文本识别方法虽然取得了很大的进步,但大多集中于拉丁字符,而针对汉字的识别方法却很少。汉字和拉丁字符之间有两个主要区别。首先,汉字的数量远大于拉丁字符。其次,汉字的宽度随字符的不同而变化,而拉丁字符的宽度是稳定的。本文提出一种多宽度文本识别方法来解决中文文本识别中的两个挑战。引入多特征提取器模块来获取中文任务中字符的多个特征。此外,采用基于注意力的融合模块来动态融合多个特征。我们在一个流行的中国数据集上进行了实验,结果表明我们的模型优于其他最先进的(SOTA)基线。
Recognizing texts plays an important role in optical characters recognition (OCR). Although the previous text recognition methods have made great progress, most of them focus on Latin characters, while few are about Chinese characters. There are two main differences between Chinese characters and Latin characters. First, the amount of Chinese characters is much larger than that of Latin characters. Second, the width of Chinese characters varies from character, while that of Latin characters is stable. This paper proposes a multi-width text recognition method to solve the two challenges in Chinese text recognition. A multiple feature extractor module is introduced to obtain multiple features of characters in a Chinese task. Besides, an attention-based fusion module is employed to dynamically fuse the multiple features. We conduct experiments on a popular Chinese dataset, and the results demonstrate that our model is superior to other state-of-the-art (SOTA) baselines.