Improving Low Resource Turkish Speech Recognition with Data Augmentation and TTS
Improving Low Resource Turkish Speech Recognition with Data Augmentation and TTS
复制标题
通过数据增强和 TTS 改进低资源土耳其语语音识别
DOI:
10.1109/ssd.2019.8893184
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
H. Yalcin
中科院分区:
文献类型:
--
作者:
Ramazan Gokay;H. Yalcin
One of the major problems faced by speech recognition researchers is the lack of data. In this paper, our objective is to compare alternative solutions to lack of data. Some experiments are conducted with very limited training data to see the effects of data augmentation and speech synthesis on speech recognition. Speed and volume perturbations are applied in this study. Besides data augmentation, synthetic speech is generated by using two different speech synthesis methods. In first speech synthesis approach, Google Translate Text to Speech (gTTS) is used as speech synthesizer. In second speech synthesis approach, an end-to-end Turkish TTS system is trained by us. Finally, we examined the effects of all these alternative methods on speech recognition for low resource languages. Our results demonstrate that some data augmentation or speech synthesis techniques work well to improve speech recognition for low resource languages. In this study, 14.8% relative Word Error Ratio (WER) improvement is obtained by using combination of augmented and synthetic data.
DOI:
--
发表时间:
2015-12
期刊:
--
影响因子:
--
作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;
通讯作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;