A New Corpus of Elderly Japanese Speech for Acoustic Modeling, and a Preliminary Investigation of Dialect-Dependent Speech Recognition

A New Corpus of Elderly Japanese Speech for Acoustic Modeling, and a Preliminary Investigation of Dialect-Dependent Speech Recognition
复制标题

DOI:
10.1109/o-cocosda46868.2019.9041216
复制
发表时间:
2019-10
期刊:
2019 22nd Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)
影响因子:
--
通讯作者:
Meiko Fukuda;Ryota Nishimura;H. Nishizaki;Y. Iribe;N. Kitaoka
Meiko Fukuda;Ryota Nishimura;H. Nishizaki;Y. Iribe;N. Kitaoka
中科院分区:
其他
文献类型:
--
作者:
Meiko Fukuda;Ryota Nishimura;H. Nishizaki;Y. Iribe;N. Kitaoka

文献摘要

相似文献

我们已经构建了一个新的语音数据语料库,包括221名日本老年人(平均年龄:79.2)的话语,目的是提高老年人的自动语音识别(ASR)的准确性。ASR对于视力受损或手部运动受限的人(包括老年人)来说是一种有益的方式。然而,语音识别系统使用标准的识别模型,特别是声学模型,一直无法达到令人满意的性能为老年人。因此,创建更准确的声学模型的语音老年用户是必不可少的,以提高语音识别的老年人。使用我们的新语料库,其中包括生活在日本三个地区的老年人的语音,我们进行了语音识别实验,使用各种DNN-HNN声学模型。作为我们声学模型的训练数据,我们研究了标准成人日语语音语料库(JNAS),老年人语音语料库(S-JNAS)或自发语音语料库(CSJ)是否最合适,以及是否适应每个地区的方言提高识别结果。我们将三个声学模型中的每一个都调整到我们所有的语音数据中,然后使用每个区域的语音重新调整它们。在没有自适应的情况下,当使用S-JNAS训练的声学模型时获得了最好的识别结果(总语料:21.85%的单词错误率)。然而,在我们的声学模型适应整个语料库之后,CSJ训练的模型达到了最低的WER(整个语料库:17.42%)。此外,在重新适应每个区域方言,CSJ训练的声学模型与适应区域语音数据显示出提高识别率的趋势。我们计划从日本各地收集更多的话语,以便我们的语料库可以用作日语老年人语音识别的关键资源。我们也希望能够进一步提高老年人语音识别的性能。
We have constructed a new speech data corpus consisting of the utterances of 221 elderly Japanese people (average age: 79.2) with the aim of improving the accuracy of automatic speech recognition (ASR) for the elderly. ASR is a beneficial modality for people with impaired vision or limited hand movement, including the elderly. However, speech recognition systems using standard recognition models, especially acoustic models, have been unable to achieve satisfactory performance for the elderly. Thus, creating more accurate acoustic models of the speech of elderly users is essential for improving speech recognition for the elderly. Using our new corpus, which includes the speech of elderly people living in three regions of Japan, we conducted speech recognition experiments using a variety of DNN-HNN acoustic models. As training data for our acoustic models, we examined whether a standard adult Japanese speech corpus (JNAS), an elderly speech corpus (S-JNAS) or a spontaneous speech corpus (CSJ) was most suitable, and whether or not adaptation to the dialect of each region improved recognition results. We adapted each of our three acoustic models to all of our speech data, and then re-adapt them using speech from each region. Without adaptation, the best recognition results were obtained when using the S-JNAS trained acoustic models (total corpus: 21.85% Word Error Rate). However, after adaptation of our acoustic models to our entire corpus, the CSJ trained models achieved the lowest WERs (entire corpus: 17.42%). Moreover, after readaptation to each regional dialect, the CSJ trained acoustic model with adaptation to regional speech data showed tendencies of improved recognition rates. We plan to collect more utterances from all over Japan, so that our corpus can be used as a key resource for elderly speech recognition in Japanese. We also hope to achieve further improvement in recognition performance for elderly speech.