High Performance Urdu and Arabic Video Text Recognition Using Convolutional Recurrent Neural Networks

High Performance Urdu and Arabic Video Text Recognition Using Convolutional Recurrent Neural Networks
复制标题

使用卷积循环神经网络的高性能乌尔都语和阿拉伯语视频文本识别

DOI:
10.1007/978-3-030-86198-8_24
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
F. Shafait
F. Shafait
中科院分区:
--
文献类型:
--
作者:
A. Rehman;A. Ul;F. Shafait

文献摘要

被引文献

相似文献

从视频中提取文本是文档分析领域的一个新兴研究领域。我们提出了一个简单的卷积递归神经网络对阿拉伯语和乌尔都语脚本进行文本识别。我们使用各种各样的数据增强技术来推广模型并防止过度拟合。我们还使用了一个稍微改进的损失函数,帮助模型更快地收敛。使用所提出的方法,我们取得了99.73%的CRR,88.37%的WRR和89.92%的LRR的乌尔都语Ticker文本数据集和96.82%的CRR,90.41%的WRR和76.78%的LRR的AcTiVComp20数据集。所提出的方法在这两个数据集上的表现都明显优于Google Vision API。
Text extraction from videos is an emerging research field in the document analysis community. We propose a simple Convolutional Recurrent Neural Network to perform text recognition on both Arabic and Urdu scripts. We use a large variety of data augmentation techniques to generalize the model and prevent over-fitting. We also use a slightly improved loss function that helps the model converge faster. Using the proposed method we achieved 99.73% CRR, 88.37% WRR and 89.92% LRR on the Urdu Ticker Text dataset and 96.82% CRR, 90.41% WRR and 76.78% LRR on the AcTiVComp20 dataset. The proposed method has significantly outperformed Google Vision API on both of the datasets.