Developing Automatic Speech Recognition for Scottish Gaelic

Developing Automatic Speech Recognition for Scottish Gaelic
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Lucy Evans;W. Lamb;M. Sinclair;Beatrice Alex
Lucy Evans;W. Lamb;M. Sinclair;Beatrice Alex
中科院分区:
其他
文献类型:
--
作者:
Lucy Evans;W. Lamb;M. Sinclair;Beatrice Alex

文献摘要

被引文献

相似文献

本文从有限的资源出发,讨论了我们为开发苏格兰盖尔语全自动语音识别(ASR)系统所做的努力。建立ASR技术对于记录和振兴濒危语言非常重要;它可以通过自动字幕和转录增强现有资源,提高用户的可访问性,并反过来鼓励继续使用该语言。在本文中,我们解释了在收集少数民族语言数据用于语音识别时所面临的许多困难。一种新的跨语言训练数据对齐方法被用来克服这样一个困难,通过这种方式,我们展示了大多数语言资源如何引导资源较少的语言技术的发展。我们使用Kaldi语音识别工具包开发了几个盖尔语ASR系统,并报告了最终的WER为26.30%。这比我们原来的型号改进了9.50%。
This paper discusses our efforts to develop a full automatic speech recognition (ASR) system for Scottish Gaelic, starting from a point of limited resource. Building ASR technology is important for documenting and revitalising endangered languages; it enables existing resources to be enhanced with automatic subtitles and transcriptions, improves accessibility for users, and, in turn, encourages continued use of the language. In this paper, we explain the many difficulties faced when collecting minority language data for speech recognition. A novel cross-lingual approach to the alignment of training data is used to overcome one such difficulty, and in this way we demonstrate how majority language resources can bootstrap the development of lower-resourced language technology. We use the Kaldi speech recognition toolkit to develop several Gaelic ASR systems, and report a final WER of 26.30%. This is a 9.50% improvement on our original model.