Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech Recognition

Fundamental Frequency Feature Normalization and Data Augmentation for Child Speech Recognition
复制标题

DOI:
10.1109/icassp39728.2021.9413801
复制
发表时间:
2021-02
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Gary Yeung;Ruchao Fan;A. Alwan
Gary Yeung;Ruchao Fan;A. Alwan
中科院分区:
其他
文献类型:
--
作者:
Gary Yeung;Ruchao Fan;A. Alwan

文献摘要

被引文献

相似文献

由于适合儿童年龄的教育技术的重要性,需要幼儿自动语音识别(ASR)系统。由于缺乏公开可用的幼儿语音数据,必须考虑特征提取策略,如特征归一化和数据增强,以成功训练儿童ASR系统。本研究提出了一种基于共振峰和基频之间关系的特征归一化和数据增强方法的儿童ASR新技术。特征归一化和数据增强技术都是在Mel域中通过频移实现的。这些技术在儿童阅读语音ASR任务中进行了评估。儿童ASR系统是通过采用基于blstm的声学模型训练成人语言来训练的。在OGI儿童语音语料库上进行测试时,使用归一化和数据增强的结果使相对单词错误率(WER)比基线提高了19.3%,并且所得到的儿童ASR系统达到了该语料库上目前报道的最佳WER。
Automatic speech recognition (ASR) systems for young children are needed due to the importance of age-appropriate educational technology. Because of the lack of publicly available young child speech data, feature extraction strategies such as feature normalization and data augmentation must be considered to successfully train child ASR systems. This study proposes a novel technique for child ASR using both feature normalization and data augmentation methods based on the relationship between formants and fundamental frequency (fo). Both the fo feature normalization and data augmentation techniques are implemented as a frequency shift in the Mel domain. These techniques are evaluated on a child read speech ASR task. Child ASR systems are trained by adapting a BLSTM-based acoustic model trained on adult speech. Using both fo normalization and data augmentation results in a relative word error rate (WER) improvement of 19.3% over the baseline when tested on the OGI Kids’ Speech Corpus, and the resulting child ASR system achieves the best WER currently reported on this corpus.