SanskritTagger , a stochastic lexical and POS tagger for Sanskrit

SanskritTagger , a stochastic lexical and POS tagger for Sanskrit
复制标题

SanskritTagger ,梵语的随机词汇和词性标注器

DOI:
--
复制
发表时间:
2007
期刊:
影响因子:
--
通讯作者:
Oliver Hellwig
Oliver Hellwig
中科院分区:
--
文献类型:
--
作者:
Oliver Hellwig

文献摘要

被引文献

相似文献

SanskritTagger是一个随机标记器,用于未预处理的梵文文本。标记器使用马尔可夫模型对文本进行标记,并使用隐马尔可夫模型执行词性标记。这些过程的参数是从一个人工注释的语料库中估计出来的,目前大约有1,500,000个单词。本文概述了标注过程,报告了几段梵文文本的标注结果,并描述了该程序的进一步改进。本文描述了SanskritTagger的设计和功能,这是一个标记器和词性(POS)标注器,它通过重复应用随机模型来分析“自然”,即未注释的梵文文本。这个标注器是在过去几年里开发出来的,作为一个更大的梵文文本数字化项目(cmp)的一部分。(Hellwig, 2002)),并仍处于稳步改善的状态。文章组织如下:第1节简要概述了影响标注器设计的梵文文本中发现的语言问题。第2节描述了标记器的实际实现。在第3节中,对来自不同主题领域的短文本进行了标注器的性能评估。此外,本节还描述了未来版本中可能的改进。
SanskritTagger is a stochastic tagger for unpreprocessed Sanskrit text. The tagger tokenises text with a Markov model and performs part-of-speech tagging with a Hidden Markov model. Parameters for these processes are estimated from a manually annotated corpus of currently about 1.500.000 words. The article sketches the tagging process, reports the results of tagging a few short passages of Sanskrit text and describes further improvements of the program. The article describes design and function of SanskritTagger, a tokeniser and part-of-speech (POS) tagger, which analyses ”natural”, i.e. unannotated Sanskrit text by repeated application of stochastic models. This tagger has been developped during the last few years as part of a larger project for digitalisation of Sanskrit texts (cmp. (Hellwig, 2002)) and is still in the state of steady improvement. The article is organised as follows: Section 1 gives a short overview about linguistic problems found in Sanskrit texts which influenced the design of the tagger. Section 2 describes the actual implementation of the tagger. In section 3, the performance of the tagger is evaluated on short passages of text from different thematic areas. In addition, this section describes possible improvements in future versions.