Utilizing Language Technology in the Documentation of Endangered Uralic Languages

Utilizing Language Technology in the Documentation of Endangered Uralic Languages
复制标题

利用语言技术记录濒危乌拉尔语言

DOI:
--
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
J. Wilbur
J. Wilbur
中科院分区:
--
文献类型:
--
作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur

文献摘要

被引文献

相似文献

本文介绍了Pite Saami、Kola Saami和Izhva Komi语言文献项目正在进行的工作,所有这些项目都记录了新的口语数据,收集了现有的录音,并对这些多媒体数据进行了注释,以便为今后研究濒危和未充分描述的乌拉尔语群体提供全面的语言语料库。在语言文档中应用语言技术有助于我们创建更系统地注释的语料库,而不是折衷的数据集合。具体来说,我们描述了一个脚本,提供不同的形态句法分析模块之间的互动有限状态传感器和ELAN,图形用户界面工具,注释和呈现多模态语料库。最终,在我们的项目中创建的口语语料库将是有用的,在这些语言在未来的科学意义上的定量调查。* 作者姓名按字母顺序排列。北方欧洲语言技术杂志,2016年,第4卷,第3条,第29-47页DOI 10.3384/nejlt.2000-1533.1643
The paper describes work-in-progress by the Pite Saami, Kola Saami and Izhva Komi language documentation projects, all of which record new spoken language data, digitize available recordings and annotate these multimedia data in order to provide comprehensive language corpora as databases for future research on and for endangered – and under-described – Uralic speech communities. Applying language technology in language documentation helps us to create more systematically annotated corpora, rather than eclectic data collections. Specifically, we describe a script providing interactivity between different morphosyntactic analysis modules implemented as Finite State Transducers and ELAN, a Graphical User Interface tool for annotating and presenting multimodal corpora. Ultimately, the spoken corpora created in our projects will be useful for scientifically significant quantitative investigations on these languages in the future. *The order of the authors’ names is alphabetical. Northern European Journal of Language Technology, 2016, Vol. 4, Article 3, pp 29–47 DOI 10.3384/nejlt.2000-1533.1643