Speech-Driven Facial Animation Using Polynomial Fusion of Features

Speech-Driven Facial Animation Using Polynomial Fusion of Features
复制标题

DOI:
10.1109/icassp40776.2020.9054469
复制
发表时间:
2019-12
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Triantafyllos Kefalas;Konstantinos Vougioukas;Yannis Panagakis;Stavros Petridis;Jean Kossaifi;M. Pantic
Triantafyllos Kefalas;Konstantinos Vougioukas;Yannis Panagakis;Stavros Petridis;Jean Kossaifi;M. Pantic
中科院分区:
其他
文献类型:
--
作者:
Triantafyllos Kefalas;Konstantinos Vougioukas;Yannis Panagakis;Stavros Petridis;Jean Kossaifi;M. Pantic

文献摘要

被引文献

相似文献

语音驱动的面部动画涉及使用语音信号来生成说话面部的逼真视频。最近的面部合成深度学习方法依赖于提取低维表示并将它们连接起来,然后对连接的向量进行解码。这只考虑了特征的一阶相互作用,而忽略了高阶相互作用。在本文中,我们提出了一个多项式融合层,模型的联合表示的编码由一个高阶多项式,与张量分解建模的参数。我们证明了这种方法的适用性,通过实验生成的视频评价的一系列指标的视频质量,视听同步和生成的眨眼。
Speech-driven facial animation involves using a speech signal to generate realistic videos of talking faces. Recent deep learning approaches to facial synthesis rely on extracting low-dimensional representations and concatenating them, followed by a decoding step of the concatenated vector. This accounts for only first-order interactions of the features and ignores higher-order interactions. In this paper we propose a polynomial fusion layer that models the joint representation of the encodings by a higher-order polynomial, with the parameters modelled by a tensor decomposition. We demonstrate the suitability of this approach through experiments on generated videos evaluated on a range of metrics on video quality, audiovisual synchronisation and generation of blinks.