Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yolóxochitl Mixtec

Leveraging End-to-End ASR for Endangered Language Documentation: An Empirical Study on Yolóxochitl Mixtec
复制标题

DOI:
10.18653/v1/2021.eacl-main.96
复制
发表时间:
2021-01
期刊:
--
影响因子:
--
通讯作者:
Jiatong Shi;Jiatong Shi. Jonathan D. Amith;Rey Castillo Garc'ia;Esteban Guadalupe Sierra;Kevin Duh;
Jiatong Shi;Jiatong Shi. Jonathan D. Amith;Rey Castillo Garc'ia;Esteban Guadalupe Sierra;Kevin Duh;
中科院分区:
其他
文献类型:
--
作者:
Jiatong Shi;Jiatong Shi. Jonathan D. Amith;Rey Castillo Garc'ia;Esteban Guadalupe Sierra;Kevin Duh;

文献摘要

相似文献

由于缺乏有效的人类转录员(即转录员短缺)而造成的“转录瓶颈”是濒危语言(EL)文献的主要挑战之一。自动语音识别(ASR)被认为是克服这些瓶颈的工具。根据这一建议,我们研究了端到端ASR的EL文档的有效性,与隐马尔可夫模型ASR系统不同,它避开了语言资源,而是更依赖于大数据设置。我们开源了一个Yoloxóchitl Mixtec EL语料库。首先,我们回顾了构建端到端ASR系统的方法,该方法可以被ASR社区复制。然后,我们提出了一个新手转录纠正任务,并演示了ASR系统和新手转录员如何协同工作来改进EL文档。我们相信这种组合方法将缓解转录瓶颈和转录员短缺,阻碍EL文档。
“Transcription bottlenecks”, created by a shortage of effective human transcribers (i.e., transcriber shortage), are one of the main challenges to endangered language (EL) documentation. Automatic speech recognition (ASR) has been suggested as a tool to overcome such bottlenecks. Following this suggestion, we investigated the effectiveness for EL documentation of end-to-end ASR, which unlike Hidden Markov Model ASR systems, eschews linguistic resources but is instead more dependent on large-data settings. We open source a Yoloxóchitl Mixtec EL corpus. First, we review our method in building an end-to-end ASR system in a way that would be reproducible by the ASR community. We then propose a novice transcription correction task and demonstrate how ASR systems and novice transcribers can work together to improve EL documentation. We believe this combinatory methodology would mitigate the transcription bottleneck and transcriber shortage that hinders EL documentation.