Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation

Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation
复制标题

使用自监督预训练和数据增强增强直接语音到语音翻译

DOI:
10.48550/arxiv.2204.02967
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Ann Lee
Ann Lee
中科院分区:
--
文献类型:
--
作者:
Sravya Popuri;Peng;Changhan Wang;J. Pino;Yossi Adi;Jiatao Gu;Wei;Ann Lee

文献摘要

被引文献

相似文献

直接语音到语音翻译(S2ST)模型存在数据稀缺问题,因为与由自动语音识别(ASR)、机器翻译(MT)和文本到语音(TTS)合成组成的常规级联系统可用的数据量相比,几乎没有并行的S2ST数据。在这项工作中,我们探索了无标签语音数据的自我监督预训练和数据扩充来解决这个问题。我们利用最近提出的语音到单元翻译(S2UT)框架,将目标语音编码成离散表示,并通过研究语音编码器和离散单元译码的预训练,将适用于语音到文本翻译(S2T)的预训练和高效的部分微调技术转移到S2UT域。我们在西班牙语-英语翻译上的实验表明,与多任务学习相比,自我监督预训练的模型性能得到了持续的提高,BLEU的平均收益为6.6%-12.1%,并且可以进一步与应用机器翻译的数据扩充技术相结合来创建弱监督训练数据。音频样本可在以下网址获得:https://facebookresearch.github.io/speech_translation/enhanced_direct_s2st_units/index.html。
Direct speech-to-speech translation (S2ST) models suffer from data scarcity issues as there exists little parallel S2ST data, compared to the amount of data available for conventional cascaded systems that consist of automatic speech recognition (ASR), machine translation (MT), and text-to-speech (TTS) synthesis. In this work, we explore self-supervised pre-training with unlabeled speech data and data augmentation to tackle this issue. We take advantage of a recently proposed speech-to-unit translation (S2UT) framework that encodes target speech into discrete representations, and transfer pre-training and efficient partial finetuning techniques that work well for speech-to-text translation (S2T) to the S2UT domain by studying both speech encoder and discrete unit decoder pre-training. Our experiments on Spanish-English translation show that self-supervised pre-training consistently improves model performance compared with multitask learning with an average 6.6-12.1 BLEU gain, and it can be further combined with data augmentation techniques that apply MT to create weakly supervised training data. Audio samples are available at: https://facebookresearch.github.io/speech_translation/enhanced_direct_s2st_units/index.html .