Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation

Speech Segmentation Optimization using Segmented Bilingual Speech Corpus for End-to-end Speech Translation
复制标题

DOI:
10.48550/arxiv.2203.15479
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Ryo Fukuda;Katsuhito Sudoh;Satoshi Nakamura
Ryo Fukuda;Katsuhito Sudoh;Satoshi Nakamura
中科院分区:
其他
文献类型:
--
作者:
Ryo Fukuda;Katsuhito Sudoh;Satoshi Nakamura

文献摘要

被引文献

相似文献

语音分割是语音翻译的基础,语音分割是将长语音分割成短片段的过程。流行的VAD工具,如WebRTC VAD,通常依赖于基于暂停的分段。不幸的是,讲话中的停顿不一定与句子边界匹配,并且句子可以通过VAD很难检测到的非常短的停顿来连接。在这项研究中,我们提出了一种使用分割后的双语语料库训练的二进制分类模型的语音分割方法。我们还提出了一种将VAD和上述语音分割方法相结合的混合方法。实验结果表明,该方法比传统的分割方法更适用于级联和端到端的ST段分割。这种混合方法进一步提高了翻译性能。
Speech segmentation, which splits long speech into short segments, is essential for speech translation (ST). Popular VAD tools like WebRTC VAD have generally relied on pause-based segmentation. Unfortunately, pauses in speech do not necessarily match sentence boundaries, and sentences can be connected by a very short pause that is difficult to detect by VAD. In this study, we propose a speech segmentation method using a binary classification model trained using a segmented bilingual speech corpus. We also propose a hybrid method that combines VAD and the above speech segmentation method. Experimental results revealed that the proposed method is more suitable for cascade and end-to-end ST systems than conventional segmentation methods. The hybrid approach further improved the translation performance.