TCD-TIMIT: An Audio-Visual Corpus of Continuous Speech

TCD-TIMIT: An Audio-Visual Corpus of Continuous Speech
复制标题

DOI:
10.1109/tmm.2015.2407694
复制
发表时间:
2015-05-01
影响因子:
7.3
通讯作者:
Gillen, Eoin
Gillen, Eoin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Harte, Naomi;Gillen, Eoin

文献摘要

被引文献

相似文献

自动视听语音识别目前在重大进展方面落后于纯音频语音识别。研究人员经常提到的原因之一是缺乏合适的研究语料库。本文详细介绍了一个新的语料库的建立,该语料库是为连续视听语音识别研究设计的。TCD-TIMIT包括62位发言者的高质量音频和视频片段,总共阅读了6913个语音丰富的句子。其中三名说话者是经过专业训练的口红说话者,录制下来是为了测试这样的假设,即在自动视觉语音识别系统中,口红说话者可能比普通说话者更有优势。视频片段是从两个角度录制的:笔直向前和在。这份文件概述了镜头的录制,以及为每句话制作视频和音频剪辑所需的后处理。报告了音频、视频和联合视听基线实验。在唇语者和非唇语者的数据上分别进行了实验,并对结果进行了比较。非口红说话者的视觉和视听基线结果总体较低。在口红音箱上的结果被发现要高得多。希望作为一个公开可用的数据库,TCD-TIMIT现在将有助于进一步提高视听语音识别研究的水平。
Automatic audio-visual speech recognition currently lags behind its audio-only counterpart in terms of major progress. One of the reasons commonly cited by researchers is the scarcity of suitable research corpora. This paper details the creation of a new corpus designed for continuous audio-visual speech recognition research. TCD-TIMIT consists of high-quality audio and video footage of 62 speakers reading a total of 6913 phonetically rich sentences. Three of the speakers are professionally-trained lipspeakers, recorded to test the hypothesis that lipspeakers may have an advantage over regular speakers in automatic visual speech recognition systems. Video footage was recorded from two angles: straight on, and at. The paper outlines the recording of footage, and the required post-processing to yield video and audio clips for each sentence. Audio, visual, and joint audio-visual baseline experiments are reported. Separate experiments were run on the lipspeaker and non-lipspeaker data, and the results compared. Visual and audio-visual baseline results on the non-lipspeakers were low overall. Results on the lipspeakers were found to be significantly higher. It is hoped that as a publicly available database, TCD-TIMIT will now help further state of the art in audio-visual speech recognition research.