The IMP historical Slovene language resources

The IMP historical Slovene language resources
复制标题

IMP历史斯洛文尼亚语言资源

DOI:
--
复制
发表时间:
2015
影响因子:
2.7
通讯作者:
T. Erjavec
T. Erjavec
中科院分区:
计算机科学4区
文献类型:
--
作者:
T. Erjavec

文献摘要

被引文献

相似文献

本文描述了几个项目的综合结果,这些项目构成了印刷历史斯洛文尼亚语的基本语言资源基础设施。IMP语言资源包括一个数字图书馆、一个带注释的语料库和一个词典,它们按照文本编码倡议指南相互链接并统一编码。图书馆拥有大约650册(大部分是完整的书),包括45,000页的传真,以及手工修改和结构化的抄本。手工注释的语料库有30万个标记,每个单词都标有其现代词形、引理、词性,如果是古单词,则标有其最接近的当代同义词。这些信息被提取到词典中,它还涵盖了一个扩展的目标注释语料库,产生了20,000个词(其中4,000个古词),50,000个现代词形式和70,000个已证实的形式。我们还开发了一个程序,对历史上的斯洛文尼亚语进行现代化、标记和归纳,并对数字图书馆进行注释,产生了一个自动注释的1500万字的语料库。为了服务于人文学科,数字图书馆和词典可以在网上阅读和浏览,语料库可以通过一个协和舞者。对于语言技术的研究和开发,这些资源可以在知识共享署名许可下以源TEI XML的形式获得。本文介绍了IMP资源(可从http://nl.ijs.si/imp/获得)的编制、编码和传播过程,并对未来的研究方向进行了总结。
The paper describes the combined results of several projects which constitute a basic language resource infrastructure for printed historical Slovene. The IMP language resources consist of a digital library, an annotated corpus and a lexicon, which are interlinked and uniformly encoded following the Text Encoding Initiative Guidelines. The library holds about 650 units (mostly complete books) consisting of facsimiles with 45,000 pages as well as hand-corrected and structured transcriptions. The hand-annotated corpus has 300,000 tokens, where each word is tagged with its modernised word form, lemma, part-of-speech and, in cases of archaic words, its nearest contemporary equivalents. This information was extracted into the lexicon, which also covers an extended target-annotated corpus, resulting in 20,000 lemmas (of these 4,000 archaic) with 50,000 modern word forms and 70,000 attested forms. We have also developed a program to modernise, tag and lemmatise historical Slovene, and annotated the digital library with it, producing an automatically annotated corpus of 15 million words. To serve the humanities, the digital library and lexicon are available for reading and browsing on the web and the corpora via a concordancer. For language technology research and development the resources are available in source TEI XML under the Creative Commons Attribution licence. The paper presents the IMP resources, available from http://nl.ijs.si/imp/, the process of their compilation, encoding and dissemination, and concludes with directions for future research.