High Performance Urdu and Arabic Video Text Recognition Using Convolutional Recurrent Neural Networks
High Performance Urdu and Arabic Video Text Recognition Using Convolutional Recurrent Neural Networks
复制标题
使用卷积循环神经网络的高性能乌尔都语和阿拉伯语视频文本识别
DOI:
10.1007/978-3-030-86198-8_24
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
F. Shafait
中科院分区:
文献类型:
--
作者:
A. Rehman;A. Ul;F. Shafait
Text extraction from videos is an emerging research field in the document analysis community. We propose a simple Convolutional Recurrent Neural Network to perform text recognition on both Arabic and Urdu scripts. We use a large variety of data augmentation techniques to generalize the model and prevent over-fitting. We also use a slightly improved loss function that helps the model converge faster. Using the proposed method we achieved 99.73% CRR, 88.37% WRR and 89.92% LRR on the Urdu Ticker Text dataset and 96.82% CRR, 90.41% WRR and 76.78% LRR on the AcTiVComp20 dataset. The proposed method has significantly outperformed Google Vision API on both of the datasets.