Improvement of Elderly Speech Recognition Using Gammatone Filterbank Adaptation

Improvement of Elderly Speech Recognition Using Gammatone Filterbank Adaptation
复制标题

使用 Gammatone 滤波器组自适应改进老年人语音识别

DOI:
10.1109/gcce53005.2021.9622086
复制
发表时间:
2021
期刊:
Proceedings of 2020 IEEE 10th Global Conference on Consumer Electronics (GCCE)
影响因子:
--
通讯作者:
Seiichi Nakagawa
Seiichi Nakagawa
中科院分区:
--
文献类型:
--
作者:
Kazumasa Yamamoto;Akinori Ishiki;Seiichi Nakagawa

文献摘要

相似文献

最近,通过引入深度学习,语音识别的准确性得到了显着提高。然而,对于老年人/高龄老人的语音识别的准确性仍然不足,并且针对他们的口语对话系统还没有投入实际使用。在这项研究中,GtFDNN-HMM,它使用了一个gammatone滤波器组的特征提取层的DNN,被用作声学模型,和扬声器/老年人的适应,以提高语音识别的准确性为老/最老的老人。GtFDNN的自适应只应用了gammatone滤波器组的参数,因此可以实现少量语音的说话人/老年人自适应,适用于难以采集大量语音的老年人/高龄老人的语音识别。实验结果表明,该模型能有效地提高老年人语音识别的性能。
Recently, the accuracy of speech recognition has been remarkably improved by the introduction of deep learning. However, the accuracy of speech recognition for old-old/oldest-old people is still insufficient and a spoken dialogue system for them has not been put into practical use. In this study, GtFDNN-HMM, which uses a gammatone filterbank for the feature extraction layer in the DNN, is used as an acoustic model, and speaker/elderly adaptation is performed to improve the accuracy of speech recognition for the old-old/oldest-old people. Only the parameters of the gammatone filterbank are applied for adaptation of GtFDNN, so speaker/elderly adaptation with a small amount of speech is possible, and it is suitable for speech recognition for old-old/oldest-old people, where it is difficult to collect a large amount of speech. From the result of experiments, the speech recognition performance was improved by using the elderly speech adaptation model.