One Size Does Not Fit All in Resource-Constrained ASR

One Size Does Not Fit All in Resource-Constrained ASR
复制标题

DOI:
10.21437/interspeech.2021-1970
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux
中科院分区:
其他
文献类型:
--
作者:
Ethan J. Morris;Robert Jimerson;Emily Prudhommeaux

文献摘要

被引文献

相似文献

将深度神经网络应用于自动语音识别的声学建模任务,大大降低了ASR单词错误率,使这项技术能够用于与高资源语言的智能手机和个人家庭助理进行交互。然而,开发这种口径的ASR模型需要数百或数千小时的转录语音记录,这对世界上绝大多数语言都提出了挑战。在本文中,我们研究了三种不同的架构,以前被用于ASR的语言与有限的培训资源的效用。我们在公开可用的ASR数据集上训练和测试这些系统,这些数据集是在各种条件下使用不同的语音收集策略、实践和设备产生的,用于几种类型和正字法不同的语言。虽然这些语料库在大小上相当,但我们发现没有一个ASR架构优于所有其他架构。此外,单词错误率差异很大,在某些情况下,在高资源语言通常报告的范围内。我们的研究结果表明,在为培训资源有限的语言开发ASR系统时,考虑特定语言和特定语料库的因素并尝试多种方法非常重要。
The application of deep neural networks to the task of acoustic modeling for automatic speech recognition has resulted in dramatic decreases in ASR word error rates, enabling the use of this technology for interacting with smart phones and personal home assistants in high-resource languages. Developing ASR models of this caliber, however, requires hundreds or thousands of hours of transcribed speech recordings, which presents challenges for the vast majority of the world’s languages. In this paper, we investigate the utility of three distinct architectures that have previously been used for ASR in languages with limited training resources. We train and test these systems on publicly available ASR datasets for several typologically and orthographi-cally diverse languages, which were produced under a variety of conditions using different speech collection strategies, practices, and equipment. Although these corpora are comparable in size, we find that no single ASR architecture outperforms all others. In addition, word error rates vary significantly, in some cases within the range of those typically reported for high-resource languages. Our results point to the importance of considering language-specific and corpus-specific factors and experimenting with multiple approaches when developing ASR systems for languages with limited training resources.