Ligature Segmentation for Urdu OCR

Ligature Segmentation for Urdu OCR
复制标题

乌尔都语 OCR 的连字分割

DOI:
10.1109/icdar.2013.229
复制
发表时间:
2013
期刊:
2013 12th International Conference on Document Analysis and Recognition
影响因子:
--
通讯作者:
Gurpreet Singh Lehal
Gurpreet Singh Lehal
中科院分区:
--
文献类型:
--
作者:
Gurpreet Singh Lehal

文献摘要

被引文献

相似文献

乌尔都语文字使用阿拉伯字母的超集,但使用纳斯塔利克书写风格。Nastaliq字体高度草书,对上下文敏感,从右上角到左下角对角线书写,字符堆叠,这使得处理OCR非常困难。此外,线条和单词切分不是简单的任务,因为我们经常要合并线条以及垂直重叠的单词和连字。由于字符分割的难点,大多数研究人员都采用下一个更高的单位--连字作为识别单位。连字是一个或多个字符的连通部分,包括变音符号,通常一个乌尔都语单词由1到8个连字组成。在本文中,我们提出了一种将乌尔都语文本切分成连字的方法。采用了一种混合方法,该方法使用自上而下技术进行线条分割,并使用自下而上设计将线条分割成连字。详细讨论了在连字分割过程中遇到的各种挑战,如水平重叠和折线、合并连字和变音符关联。
Urdu script uses superset of Arabic alphabet, but uses Nastaliq writing style. Nastaliq script is highly cursive, context sensitive and is written diagonally from top right to bottom left with stacking of characters, which makes it very hard to process for OCR. In addition, line and word segmentation are non-trivial tasks as we have frequently merging lines and vertically overlapping words and ligatures. Due to the challanges in character segmentation most of the researchers have taken the next higher unit, ligature, as recognition unit. A ligature is a connected component of one or more characters including diacritic marks and usually an Urdu word is composed of 1 to 8 ligatures. In this paper, we present a methodology for segmenting the Urdu text into ligatures. A hybrid approach, which uses top down technique for line segmentation and bottom up design for segmenting the line into ligatures, has been employed. The various challenges encountered during ligature segmentation such as horizontally overlapping and broken lines, merged ligatures and diacritic association have been discussed in detail.