Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation

Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation
复制标题

乌尔都语文本标记化和句子边界消歧的挑战

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
U. I. Bajwa
U. I. Bajwa
中科院分区:
--
文献类型:
--
作者:
Zobia Rehman;W. Anwar;U. I. Bajwa

文献摘要

被引文献

相似文献

乌尔都语是一种形态丰富、文字性质各异的语言。与英语等语言相比,乌尔都语文本标记化和句子边界消歧是困难的。标记化的主要障碍是词间空格的使用不当,由于没有大小写区分,使得句子边界检测成为一项困难的任务。在本文中,已经确定了关于这两种语言处理任务的一些问题。
Urdu is morphologically rich language with different nature of its characters. Urdu text tokenization and sentence boundary disambiguation is difficult as compared to the language like English. Major hurdle for tokenization is improper use of space between words, where as absence of case discrimination makes the sentence boundary detection a difficult task. In this paper some issues regarding both of these language processing tasks have been identified.