The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software

The Indigenous Languages Technology project at NRC Canada: An empowerment-oriented approach to developing language software
复制标题

加拿大 NRC 的土著语言技术项目:以赋权为导向的语言软件开发方法

DOI:
10.18653/v1/2020.coling-main.516
复制
发表时间:
2020
期刊:
Journal of Language, Identity & Education
影响因子:
--
通讯作者:
Heather Souter
Heather Souter
中科院分区:
--
文献类型:
--
作者:
R. Kuhn;Fineen Davis;Alain Désilets;E. Joanis;Anna Kazantseva;Rebecca Knowles;Patrick Littell;Delaney Lothian;Aidan Pine;Caroline Running Wolf;E. Santos;Darlene A. Stewart;Gilles Boulianne;Vishwa Gupta;Brian Maracle Owennatékha;Akwiratékha’ Martin;Christopher Cox;M. Junker;Olivia N. Sammons;D. Torkornoo;Nathan Thanyehténhas Brinklow;Sara Child;Benoit Farley;David Huggins;Daisy Rosenblum;Heather Souter

文献摘要

被引文献

相似文献

本文调查了加拿大国家研究理事会的一个项目的第一个为期三年的阶段,该项目正在开发软件,以帮助加拿大土著社区保护他们的语言并扩大其使用。该项目的目的是在增强权能的范式内开展工作,其中与社区合作并实现其目标是核心。由于我们开发的许多技术都是为了响应社区的需求,因此该项目最终成为一个不同子项目的集合,包括创建一个复杂的框架,用于为高度屈折的多合成语言构建动词变位器(例如易洛魁语系的Kanyen'kéha),发布可能是最大的可用语料库的句子在一个多合成语言(Inuktut)与英语句子和实验与机器翻译(MT)对齐在此语料库上训练的系统,基于自动语音识别(ASR)的免费在线服务,以缓解土著语言(和其他语言)语音录音的转录瓶颈,用于实现文本预测和土著语言有声读物的软件,以及其他几个子项目。
This paper surveys the first, three-year phase of a project at the National Research Council of Canada that is developing software to assist Indigenous communities in Canada in preserving their languages and extending their use. The project aimed to work within the empowerment paradigm, where collaboration with communities and fulfillment of their goals is central. Since many of the technologies we developed were in response to community needs, the project ended up as a collection of diverse subprojects, including the creation of a sophisticated framework for building verb conjugators for highly inflectional polysynthetic languages (such as Kanyen’kéha, in the Iroquoian language family), release of what is probably the largest available corpus of sentences in a polysynthetic language (Inuktut) aligned with English sentences and experiments with machine translation (MT) systems trained on this corpus, free online services based on automatic speech recognition (ASR) for easing the transcription bottleneck for recordings of speech in Indigenous languages (and other languages), software for implementing text prediction and read-along audiobooks for Indigenous languages, and several other subprojects.