Instant Annotations – Applying NLP Methods to the Annotation of Spoken Language Documentation Corpora

Instant Annotations – Applying NLP Methods to the Annotation of Spoken Language Documentation Corpora
复制标题

DOI:
10.18653/v1/w17-0604
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur
中科院分区:
其他
文献类型:
--
作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur

文献摘要

被引文献

相似文献

本文件介绍了Pite Saami、Kola Saami和Izhva Komi语言文件项目正在进行的工作,所有这些项目都使用类似的数据和技术框架,在弗赖堡进行,并与汉堡、Syktyvkar、特罗姆瑟和乌普萨拉合作进行。我们的项目工作在濒危语言文件框架和记录新的口语数据,检索可用的录音和注释这些多媒体数据,以提供全面的语言语料库作为数据库,为未来的研究和濒危和未充分描述的乌拉尔语音社区。在语言文档中应用NLP方法-特别是基于规则的形态和句法分析器-可以帮助我们创建更系统的注释语料库,而不是折衷的数据集合。我们提出了一个循序渐进的方法,通过使用和改进真正的计算方法来达到更高级别的注释。最终,我们的项目所创建的口语语料库将有助于未来对这些语言进行科学上有意义的定量研究。
Thepaper describes work-in-progress by the Pite Saami, Kola Saami and Izhva Komi language documentation projects, all of which use similar data and technical frameworks and are carried out in Freiburg and in collaboration with Hamburg, Syktyvkar, Tromsø and Uppsala. Our projects work in the endangered language documentation framework and record new spoken language data, digitize available recordings and annotate these multimedia data in order to provide comprehensive language corpora as databases for future research on and for endangered and under-described Uralic speech communities. Applying NLP methods in language documentation – specifically rule-based morphological and syntactic analyzers – helps us to create more systematically annotated corpora, rather than eclectic data collections. We propose a step-by-step approach to reach higherlevel annotations by using and improving truly computational methods. Ultimately, the spoken corpora created by our projects will be useful for scientifically significant quantitative investigations on these languages in the future.