A novel holistic unconstrained handwritten urdu recognition system using convolutional neural networks

A novel holistic unconstrained handwritten urdu recognition system using convolutional neural networks
复制标题

DOI:
10.1007/s10032-022-00414-7
复制
发表时间:
2022-10-07
影响因子:
2.3
通讯作者:
Khursheed, Farida
Khursheed, Farida
中科院分区:
计算机科学4区
文献类型:
--
作者:
Ganai, Aejaz Farooq;Khursheed, Farida

文献摘要

被引文献

相似文献

由于缺乏标准的手写乌尔都语数据集、不同乌尔都语作家的写作风格差异巨大、与连字相关的变音符号位置不规则、一些乌尔都语书写字符形状相似以及缺乏有效的学习和训练技术,手写乌尔都语识别迄今为止被探索得最少。很少有研究人员提出手写乌尔都语数据集,其中只有乌尔都语 Nastaliq 手写数据集(UNHD)是公开的。 UNHD 只包含最多五个字符的连字,并不涵盖整个乌尔都语连字语料库。因此,我们为“乌尔都语手写连字数据集”提出了一个新颖的综合手写乌尔都语数据集 UHLD:-它由最多七个字符长度的连字组成,涵盖了乌尔都语语言的大部分连字语料库。 UHLD 由男女书写,与年龄、纸张颜色、纸张类型(空白或直纹)、墨水颜色、笔类型无关。我们提出了一种无约束的手写乌尔都语识别系统,可以识别最多六个字符的手写乌尔都语连字。这里还提出了一种新的鲁棒算法,能够在大型乌尔都语数据集上将完整的连字分为主要部分和次要部分,准确率达到 98%。我们提出的整体手写乌尔都语识别系统可确保独立识别单词/连字的主要和次要组成部分。所提出的识别技术具有变换不变性和计算效率,并且对于 UHLD 和 UNHD 的识别率分别达到 97% 和 93%。
Handwritten Urdu recognition has been the least explored to date due to unavailability of a standard hand-written Urdu dataset, huge variation among writing styles of different Urdu writers, irregular positioning of diacritics associated with ligatures, similarity in shape of some Urdu characters in writing, and unavailability of an efficient learning and training technique. Few researchers have proposed the handwritten Urdu datasets among which only Urdu Nastaliq handwritten dataset (UNHD) is publicly available. The UNHD contains ligatures of only up to five characters and does not cover the entire Urdu ligature corpus. Hence, we present a novel comprehensive handwritten Urdu dataset named UHLD for the 'Urdu Handwritten Ligature Dataset':-which consists of ligatures of up to seven-character length and covers most of the ligature corpus of the Urdu language. The UHLD is written by both genders independent of age of person, paper color, paper type (blank or ruled), ink color, pen type. We propose an unconstrained handwritten Urdu recognition system that can recognize handwritten Urdu ligatures with up to six characters. A new robust algorithm has also been proposed here that is able to divide a complete ligature into primary and secondary components with 98% accuracy on a large Urdu dataset. Our proposed holistic handwritten Urdu recognition system ensures independent recognition of both primary and secondary components of a word/ligature. The proposed recognition technique is transformation invariant and computationally efficient and achieves a better recognition rate of 97% for UHLD and 93% for UNHD.