Dynamic Facial Models for Video-Based Dimensional Affect Estimation

Dynamic Facial Models for Video-Based Dimensional Affect Estimation
复制标题

DOI:
10.1109/iccvw.2019.00200
复制
发表时间:
2019-10
期刊:
2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW)
影响因子:
--
通讯作者:
Siyang Song;Enrique Sánchez-Lozano;M. Tellamekala;Linlin Shen;A. Johnston;M. Valstar
Siyang Song;Enrique Sánchez-Lozano;M. Tellamekala;Linlin Shen;A. Johnston;M. Valstar
中科院分区:
其他
文献类型:
--
作者:
Siyang Song;Enrique Sánchez-Lozano;M. Tellamekala;Linlin Shen;A. Johnston;M. Valstar

文献摘要

相似文献

人脸视频的维度影响估计是一项具有挑战性的任务,主要是因为大量可能的面部显示由一组行为基元组成,包括面部肌肉动作。这种表现不仅在组成上不同,而且在时间上也不同,每一种表现都是由短期和长期特征不同的行为原语组成的。大多数现有的工作模型都依赖于复杂的分层循环模型,无法很好地捕捉短期动态。在本文中,我们提出在图像中编码这些短期的面部形状和外观动态,其中只有语义上有意义的信息被编码到动态的面部图像中。我们还提出了二进制动态面具,以从动态图像中去除“稳定像素”。这个过程允许过滤非动态信息,即只有在序列中改变的像素被保留。然后,最后提出的动态面部模型(DFM)将给定帧之前的图像序列的过滤后的面部外观和形状动态编码为三通道光栅图像。CNN-RNN架构的主要任务是对长期变化进行建模。实验表明,我们的动态人脸图像在维度影响预测任务上取得了优于标准RGB人脸图像的效果。
Dimensional affect estimation from a face video is a challenging task, mainly due to the large number of possible facial displays made up of a set of behaviour primitives including facial muscle actions. The displays vary not only in composition but also in temporal evolution, with each display composed of behaviour primitives with varying in their short and long-term characteristics. Most existing work models affect relies on complex hierarchical recurrent models unable to capture short-term dynamics well. In this paper, we propose to encode these short-term facial shape and appearance dynamics in an image, where only the semantic meaningful information is encoded into the dynamic face images. We also propose binary dynamic facial masks to remove 'stable pixels' from the dynamic images. This process allows filtering of non-dynamic information, i.e. only pixels that have changed in the sequence are retained. Then, the final proposed Dynamic Facial Model (DFM) encodes both filtered facial appearance and shape dynamics of a image sequence preceding to the given frame into a three-channel raster image. A CNN-RNN architecture is tasked with modelling primarily the long-term changes. Experiments show that our dynamic face images achieved superior performance over the standard RGB face images on dimensional affect prediction task.