Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR

Direct Simultaneous Speech-to-Text Translation Assisted by Synchronized Streaming ASR
复制标题

DOI:
10.18653/v1/2021.findings-acl.406
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Junkun Chen;Mingbo Ma;Renjie Zheng;Liang Huang
Junkun Chen;Mingbo Ma;Renjie Zheng;Liang Huang
中科院分区:
其他
文献类型:
--
作者:
Junkun Chen;Mingbo Ma;Renjie Zheng;Liang Huang

文献摘要

相似文献

同步语音到文本翻译在许多情况下都非常有用。传统的级联方法使用流ASR流水线,然后是同步MT,但存在错误传播和额外延迟。为了缓解这些问题,最近的努力试图同时将源语言直接翻译成目标文本,但由于两个独立任务的结合,这一点要困难得多。相反,我们提出了一种新的范例,兼具级联和端到端方法的优势。其核心思想是在流ASR和直接语音到文本翻译(ST)上分别使用两个独立但同步的解码器,ASR的中间结果指导ST的解码策略(但不作为输入提供给ST)。在训练时间内,我们使用多任务学习来使用一个共享编码器来联合学习这两个任务。在MuSTC数据集上进行的En-to-De和En-to-ES实验表明,我们提出的方法在相同的延迟水平下获得了更好的翻译质量。
Simultaneous speech-to-text translation is widely useful in many scenarios. The conventional cascaded approach uses a pipeline of streaming ASR followed by simultaneous MT, but suffers from error propagation and extra latency. To alleviate these issues, recent efforts attempt to directly translate the source speech into target text simultaneously, but this is much harder due to the combination of two separate tasks. We instead propose a new paradigm with the advantages of both cascaded and end-to-end approaches. The key idea is to use two separate, but synchronized, decoders on streaming ASR and direct speech-to-text translation (ST), respectively, and the intermediate results of ASR guide the decoding policy of (but is not fed as input to) ST. During training time, we use multitask learning to jointly learn these two tasks with a shared encoder. En-to-De and En-to-Es experiments on the MuSTC dataset demonstrate that our proposed technique achieves substantially better translation quality at similar levels of latency.