Video-Driven Speech Reconstruction using Generative Adversarial Networks

Video-Driven Speech Reconstruction using Generative Adversarial Networks
复制标题

DOI:
10.21437/interspeech.2019-1445
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Konstantinos Vougioukas;Pingchuan Ma;Stavros Petridis;M. Pantic
Konstantinos Vougioukas;Pingchuan Ma;Stavros Petridis;M. Pantic
中科院分区:
其他
文献类型:
--
作者:
Konstantinos Vougioukas;Pingchuan Ma;Stavros Petridis;M. Pantic

文献摘要

相似文献

言语是一种依赖于视听信息的交流手段。缺少一种模态通常会导致信息的混淆或误解。在本文中,我们提出了一个端到端时间模型,能够直接从无声视频合成音频,而无需转换中间特征。我们提出的基于gan的方法能够产生自然的声音,可理解的语音,并与视频同步。我们的模型在GRID数据集上对演讲者依赖和演讲者独立场景的性能进行了评估。据我们所知,这是第一个将视频直接映射到原始音频的方法,也是第一个在之前未见过的扬声器上测试时产生可理解语音的方法。我们对合成音频的评价不仅基于音质,还基于语音的准确性。
Speech is a means of communication which relies on both audio and visual information. The absence of one modality can often lead to confusion or misinterpretation of information. In this paper we present an end-to-end temporal model capable of directly synthesising audio from silent video, without needing to transform to-and-from intermediate features. Our proposed approach, based on GANs is capable of producing natural sounding, intelligible speech which is synchronised with the video. The performance of our model is evaluated on the GRID dataset for both speaker dependent and speaker independent scenarios. To the best of our knowledge this is the first method that maps video directly to raw audio and the first to produce intelligible speech when tested on previously unseen speakers. We evaluate the synthesised audio not only based on the sound quality but also on the accuracy of the spoken words.