Developing a Shared Task for Speech Processing on Endangered Languages

Developing a Shared Task for Speech Processing on Endangered Languages
复制标题

DOI:
10.33011/computel.v1i.967
复制
发表时间:
2021-03
期刊:
Proceedings of the Workshop on Computational Methods for Endangered Languages
影响因子:
--
通讯作者:
Gina-Anne Levow;Emily Ahn;Emily M. Bender
Gina-Anne Levow;Emily Ahn;Emily M. Bender
中科院分区:
其他
文献类型:
--
作者:
Gina-Anne Levow;Emily Ahn;Emily M. Bender

文献摘要

被引文献

相似文献

语音和语言处理方面的进步使得能够创建原则上可以加速语言文档化过程的应用程序,因为语音社区和语言学家正在进行紧急的语言文档化和回收项目。然而,由于资源需求限制了这些新技术的广泛适用性,这种系统尚未对语言文档产生重大影响。我们的目标是利用共享任务的框架,将技术研究社区集中在解决语言文档中关键痛点的任务上。在这里,我们提出了这些新的共享任务的实施的初始步骤,通过创建濒危语言库和基线系统的数据集,以执行这些音频记录的分割和扬声器标记的重要的使能步骤的文件过程。本文通过一个用例来激发这些任务,描述数据集管理和基线系统,并展示这些数据的结果。然后,我们强调的挑战和道德考虑,在开发这些语音处理工具和任务,以支持濒危语言文档。
Advances in speech and language processing have enabled the creation of applications that could, in principle, accelerate the process of language documentation, as speech communities and linguists work on urgent language documentation and reclamation projects. However, such systems have yet to make a significant impact on language documentation, as resource requirements limit the broad applicability of these new techniques. We aim to exploit the framework of shared tasks to focus the technology research community on tasks which address key pain points in language documentation. Here we present initial steps in the implementation of these new shared tasks, through the creation of data sets drawn from endangered language repositories and baseline systems to perform segmentation and speaker labeling of these audio recordings—important enabling steps in the documentation process. This paper motivates these tasks with a use case, describes data set curation and baseline systems, and presents results on this data. We then highlight the challenges and ethical considerations in developing these speech processing tools and tasks to support endangered language documentation.