Developing the Old Tibetan Treebank

Developing the Old Tibetan Treebank
复制标题

DOI:
10.26615/978-954-452-056-4_035
复制
发表时间:
2019-10
期刊:
影响因子:
3.9
通讯作者:
Christian Faggionato;M. Meelen
Christian Faggionato;M. Meelen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Christian Faggionato;M. Meelen

文献摘要

被引文献

相似文献

本文介绍了一个切分、词性标注、组块分析的古藏文语料库的开发过程。作为一种资源极其匮乏的语言,古藏语在开发可搜索树库的每一步都面临着重大问题。然而,我们证明,一个精心开发的,半监督的方法,优化和扩展现有的工具,为古典藏文,以及创建特定的古藏文可以解决这些问题。因此,我们也提出了第一个非常藏文树库在各种格式,以促进在自然语言处理,历史语言学和藏学领域的研究。
This paper presents a full procedure for the development of a segmented, POS-tagged and chunkparsed corpus of Old Tibetan. As an extremely low-resource language, Old Tibetan poses non-trivial problems in every step towards the development of a searchable treebank. We demonstrate, however, that a carefully developed, semisupervised method of optimising and extending existing tools for Classical Tibetan, as well as creating specific ones for Old Tibetan can address these issues. We thus also present the first very Tibetan Treebank in a variety of formats to facilitate research in the fields of NLP, historical linguistics and Tibetan Studies.