SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set

SPEECH-COCO: 600k Visually Grounded Spoken Captions Aligned to MSCOCO Data Set
复制标题

SPEECH-COCO:与 MSCOCO 数据集对齐的 600k 视觉基础语音字幕

DOI:
--
复制
发表时间:
2017
期刊:
Interspeech
影响因子:
--
通讯作者:
O. Rosec
O. Rosec
中科院分区:
--
文献类型:
--
作者:
William N. Havard;L. Besacier;O. Rosec

文献摘要

被引文献

相似文献

本文提出了一种增强MSCOCO数据集的语音添加到图像和文本。语音字幕使用文本到语音(TTS)合成生成,导致616,767个语音字幕(超过600小时)与图像配对。不流畅和速度扰动被添加到信号中,以便听起来更自然。每个语音信号(WAV)都与一个JSON文件配对,该文件包含口语字幕中每个单词/音节/音素的精确时间码。这样的语料库可以用于语言和视觉(LaVi)任务,包括语音输入或输出,而不是文本。调查多模态学习计划的无监督语音模式发现也有可能与此语料库,所示的一个子集的语料库(10 h,10 k口语字幕)进行的初步研究。
This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images. Disfluencies and speed perturbation are added to the signal in order to sound more natural. Each speech signal (WAV) is paired with a JSON file containing exact timecode for each word/syllable/phoneme in the spoken caption. Such a corpus could be used for Language and Vision (LaVi) tasks including speech input or output instead of text. Investigating multimodal learning schemes for unsupervised speech pattern discovery is also possible with this corpus, as demonstrated by a preliminary study conducted on a subset of the corpus (10h, 10k spoken captions).