Improving Low Resource Code-switched ASR using Augmented Code-switched TTS

Improving Low Resource Code-switched ASR using Augmented Code-switched TTS
复制标题

使用增强型代码转换 TTS 改进低资源代码转换 ASR

DOI:
10.21437/interspeech.2020-2402
复制
发表时间:
2020
期刊:
1995 International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
P. Jyothi
P. Jyothi
中科院分区:
--
文献类型:
--
作者:
Y. Sharma;Basil Abraham;Karan Taneja;P. Jyothi

文献摘要

被引文献

相似文献

由于语音技术在全球多语言社区的广泛使用,为代码转换语音构建自动语音识别 (ASR) 系统最近重新受到关注。端到端 ASR 系统是一种自然的建模选择,因为它们易于使用且在单语言环境中具有卓越的性能。然而,众所周知,端到端系统需要大量标记语音。在这项工作中,我们研究了通过使用语码转换文本到语音 (TTS) 合成的数据增强来改进低资源环境中的语码转换 ASR。我们提出了两种有针对性的技术来有效利用 TTS 语音样本:1) Mixup,一种通过现有样本的线性插值创建新训练样本的现有技术,应用于 TTS 和真实语音样本;2) 一种新的损失函数,与 TTS 样本结合使用,以鼓励代码转换预测。我们报告称,在印地语-英语语码转换 ASR 任务中,使用我们提出的技术,ASR 性能显着提高,绝对字错误率 (WER) 降低了高达 5%,并且语码转换显着改善。
Building Automatic Speech Recognition (ASR) systems for code-switched speech has recently gained renewed attention due to the widespread use of speech technologies in multilingual communities worldwide. End-to-end ASR systems are a natural modeling choice due to their ease of use and superior performance in monolingual settings. However, it is well known that end-to-end systems require large amounts of labeled speech. In this work, we investigate improving code-switched ASR in low resource settings via data augmentation using code-switched text-to-speech (TTS) synthesis. We propose two targeted techniques to effectively leverage TTS speech samples: 1) Mixup, an existing technique to create new training samples via linear interpolation of existing samples, applied to TTS and real speech samples, and 2) a new loss function, used in conjunction with TTS samples, to encourage code-switched predictions. We report significant improvements in ASR performance achieving absolute word error rate (WER) reductions of up to 5%, and measurable improvement in code switching using our proposed techniques on a Hindi-English code-switched ASR task.