Urdu language processing: a survey

Urdu language processing: a survey
复制标题

DOI:
10.1007/s10462-016-9482-x
复制
发表时间:
2017-03-01
影响因子:
12
通讯作者:
Che, Dunren
Che, Dunren
中科院分区:
计算机科学2区
文献类型:
--
作者:
Daud, Ali;Khan, Wahab;Che, Dunren

文献摘要

被引文献

相似文献

与东方语言(特别是南亚语言)相比,西方语言的自然语言处理不同活动已经进行了大量的工作。西方语言被称为资源丰富的语言。核心语言资源,例如通常可以使用为西方语言开发的语料库、WordNet、词典、地名词典和相关工具。大多数南亚语言都是资源匮乏的语言,例如乌尔都语是南亚语言,是次大陆广泛使用的语言之一。由于资源稀缺,乌尔都语方面的工作还不够。本文的核心目标是对乌尔都语语言处理中存在的不同语言资源进行调查,突出乌尔都语语言处理中的不同任务,并讨论不同的最先进的可用技术。最后,本文试图详细描述最近人们对乌尔都语语言处理研究兴趣的增长和取得的进展。首先,讨论乌尔都语语言的可用数据集。提供了乌尔都语的特征、印地语和乌尔都语之间的资源共享、正字法和形态。说明了预处理活动的各个方面,例如停用词删除、变音符号删除、标准化和词干提取。讨论了对诸如分词、句子边界检测、词性标记、命名实体识别、WordNet 任务的解析和开发等任务的最新研究的回顾。此外,还研究了 ULP 对信息检索、分类和抄袭检测等应用领域的影响。最后,提出了这个新的、充满活力的研究领域的未决问题和未来方向。本文的目标是以一种能够为未来 ULP 研究活动提供平台的方式来组织 ULP 工作。
Extensive work has been done on different activities of natural language processing for Western languages as compared to its Eastern counterparts particularly South Asian Languages. Western languages are termed as resource-rich languages. Core linguistic resources e.g. corpora, WordNet, dictionaries, gazetteers and associated tools being developed for Western languages are customarily available. Most South Asian Languages are low resource languages e.g. Urdu is a South Asian Language, which is among the widely spoken languages of sub-continent. Due to resources scarcity not enough work has been conducted for Urdu. The core objective of this paper is to present a survey regarding different linguistic resources that exist for Urdu language processing, to highlight different tasks in Urdu language processing and to discuss different state of the art available techniques. Conclusively, this paper attempts to describe in detail the recent increase in interest and progress made in Urdu language processing research. Initially, the available datasets for Urdu language are discussed. Characteristic, resource sharing between Hindi and Urdu, orthography, and morphology of Urdu language are provided. The aspects of the pre-processing activities such as stop words removal, Diacritics removal, Normalization and Stemming are illustrated. A review of state of the art research for the tasks such as Tokenization, Sentence Boundary Detection, Part of Speech tagging, Named Entity Recognition, Parsing and development of WordNet tasks are discussed. In addition, impact of ULP on application areas, such as, Information Retrieval, Classification and plagiarism detection is investigated. Finally, open issues and future directions for this new and dynamic area of research are provided. The goal of this paper is to organize the ULP work in a way that it can provide a platform for ULP research activities in future.