Improving Low Resource Turkish Speech Recognition with Data Augmentation and TTS

Improving Low Resource Turkish Speech Recognition with Data Augmentation and TTS
复制标题

通过数据增强和 TTS 改进低资源土耳其语语音识别

DOI:
10.1109/ssd.2019.8893184
复制
发表时间:
2019
期刊:
2019 16th International Multi-Conference on Systems, Signals & Devices (SSD)
影响因子:
--
通讯作者:
H. Yalcin
H. Yalcin
中科院分区:
--
文献类型:
--
作者:
Ramazan Gokay;H. Yalcin

文献摘要

参考文献

被引文献

相似文献

语音识别研究人员面临的主要问题之一是缺乏数据。在本文中,我们的目标是比较替代解决方案缺乏数据。用有限的训练数据进行了一些实验,以观察数据增强和语音合成对语音识别的影响。在这项研究中,速度和体积扰动。除了数据增强,合成语音是通过使用两种不同的语音合成方法。在第一种语音合成方法中,使用Google翻译文本到语音(gTTS)作为语音合成器。在第二种语音合成方法中,我们训练了一个端到端的土耳其语TTS系统。最后,我们研究了所有这些替代方法对低资源语言语音识别的影响。我们的研究结果表明,一些数据增强或语音合成技术工作良好,以提高低资源语言的语音识别。在这项研究中,14.8%的相对字错误率(WER)的改善,通过使用增强和合成数据的组合。
One of the major problems faced by speech recognition researchers is the lack of data. In this paper, our objective is to compare alternative solutions to lack of data. Some experiments are conducted with very limited training data to see the effects of data augmentation and speech synthesis on speech recognition. Speed and volume perturbations are applied in this study. Besides data augmentation, synthetic speech is generated by using two different speech synthesis methods. In first speech synthesis approach, Google Translate Text to Speech (gTTS) is used as speech synthesizer. In second speech synthesis approach, an end-to-end Turkish TTS system is trained by us. Finally, we examined the effects of all these alternative methods on speech recognition for low resource languages. Our results demonstrate that some data augmentation or speech synthesis techniques work well to improve speech recognition for low resource languages. In this study, 14.8% relative Word Error Ratio (WER) improvement is obtained by using combination of augmented and synthetic data.
DOI: --
发表时间: 2015-12
期刊: --
影响因子: --
作者:
Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;
通讯作者: Dario Amodei;S. Ananthanarayanan;Rishita Anubhai;Jin Bai;Eric Battenberg;Carl Case;J. Casper;