Instant Annotations – Applying NLP Methods to the Annotation of Spoken Language Documentation Corpora
Instant Annotations – Applying NLP Methods to the Annotation of Spoken Language Documentation Corpora
复制标题
DOI:
10.18653/v1/w17-0604
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur
中科院分区:
文献类型:
--
作者:
C. Gerstenberger;N. Partanen;Michael Rießler;J. Wilbur
Thepaper describes work-in-progress by the Pite Saami, Kola Saami and Izhva Komi language documentation projects, all of which use similar data and technical frameworks and are carried out in Freiburg and in collaboration with Hamburg, Syktyvkar, Tromsø and Uppsala. Our projects work in the endangered language documentation framework and record new spoken language data, digitize available recordings and annotate these multimedia data in order to provide comprehensive language corpora as databases for future research on and for endangered and under-described Uralic speech communities. Applying NLP methods in language documentation – specifically rule-based morphological and syntactic analyzers – helps us to create more systematically annotated corpora, rather than eclectic data collections. We propose a step-by-step approach to reach higherlevel annotations by using and improving truly computational methods. Ultimately, the spoken corpora created by our projects will be useful for scientifically significant quantitative investigations on these languages in the future.