Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation
Challenges in Urdu Text Tokenization and Sentence Boundary Disambiguation
复制标题
乌尔都语文本标记化和句子边界消歧的挑战
DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
U. I. Bajwa
中科院分区:
文献类型:
--
作者:
Zobia Rehman;W. Anwar;U. I. Bajwa
Urdu is morphologically rich language with different nature of its characters. Urdu text tokenization and sentence boundary disambiguation is difficult as compared to the language like English. Major hurdle for tokenization is improper use of space between words, where as absence of case discrimination makes the sentence boundary detection a difficult task. In this paper some issues regarding both of these language processing tasks have been identified.