MuST-C: a Multilingual Speech Translation Corpus

MuST-C: a Multilingual Speech Translation Corpus
复制标题

DOI:
10.18653/v1/n19-1202
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Mattia Antonino Di Gangi;R. Cattoni;L. Bentivogli;Matteo Negri;Marco Turchi
Mattia Antonino Di Gangi;R. Cattoni;L. Bentivogli;Matteo Negri;Marco Turchi
中科院分区:
其他
文献类型:
--
作者:
Mattia Antonino Di Gangi;R. Cattoni;L. Bentivogli;Matteo Negri;Marco Turchi

文献摘要

被引文献

相似文献

目前口语翻译的研究面临着规模庞大且公开可用的训练语料库缺乏的问题。这个问题阻碍了神经端到端方法的采用,神经端到端方法代表了机器翻译的两个父任务:自动语音识别和机器翻译。为了填补这一空白,我们创建了一个多语言语音翻译语料库Must-C,其规模和质量将有助于从英语到8种语言的端到端语音翻译系统的培训。对于每种目标语言,Must-C包含至少385小时的英语TED演讲录音,这些录音在句子层面上与其手动翻译和翻译自动对齐。连同语料库创建方法的描述(可扩展以添加新数据并覆盖新语言),我们提供了其质量的实证验证和使用每个语言方向上最先进的方法计算的结果。
Current research on spoken language translation (SLT) has to confront with the scarcity of sizeable and publicly available training corpora. This problem hinders the adoption of neural end-to-end approaches, which represent the state of the art in the two parent tasks of SLT: automatic speech recognition and machine translation. To fill this gap, we created MuST-C, a multilingual speech translation corpus whose size and quality will facilitate the training of end-to-end systems for SLT from English into 8 languages. For each target language, MuST-C comprises at least 385 hours of audio recordings from English TED Talks, which are automatically aligned at the sentence level with their manual transcriptions and translations. Together with a description of the corpus creation methodology (scalable to add new data and cover new languages), we provide an empirical verification of its quality and SLT results computed with a state-of-the-art approach on each language direction.