Speech-Driven Facial Animation Using Polynomial Fusion of Features
Speech-Driven Facial Animation Using Polynomial Fusion of Features
复制标题
DOI:
10.1109/icassp40776.2020.9054469
复制
发表时间:
2019-12
期刊:
影响因子:
--
通讯作者:
Triantafyllos Kefalas;Konstantinos Vougioukas;Yannis Panagakis;Stavros Petridis;Jean Kossaifi;M. Pantic
中科院分区:
文献类型:
--
作者:
Triantafyllos Kefalas;Konstantinos Vougioukas;Yannis Panagakis;Stavros Petridis;Jean Kossaifi;M. Pantic
Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This accounts for only first-order interactions of the features and ignores higher-order interactions. In this paper we propose a polynomial fusion layer that models the joint representation of the encodings by a higher-order polynomial, with the parameters modelled by a tensor decomposition. We demonstrate the suitability of this approach through experiments on generated videos evaluated on a range of metrics on video quality, audiovisual synchronisation and generation of blinks.