Temporally Coherent Visual Representations for Dimensional Affect Recognition
Temporally Coherent Visual Representations for Dimensional Affect Recognition
复制标题
DOI:
10.1109/acii.2019.8925529
复制
发表时间:
2019-09
期刊:
影响因子:
--
通讯作者:
M. Tellamekala;M. Valstar
中科院分区:
文献类型:
--
作者:
M. Tellamekala;M. Valstar
The success of end-to-end supervised representation learning and regression in recent years has shifted the main focus of continuous emotion recognition approaches to learning visual representations directly from large scale labeled datasets. Supervised representation learning is highly effective for learning tasks with unambiguous labels. But annotating dimensional affect data is inherently subjective which more often than not leads to ambiguous labels. Relying only on such ambiguous labels for representation learning does not result in robust features with good generalization capacity, as we will show in this work. To address this fundamental problem, we propose to apply a constrained representation learning method that encourages the latent features to be less sensitive to ambiguous emotion annotations by using a generic representation learning prior called ‘temporal coherency’ or ‘temporal smoothness’. This approach forces the latent features to be temporally coherent, by adding a first-order temporal coherency regularization constraint to the supervised learning loss function. To assess the utility of temporally coherent representations, we trained both unconstrained and temporal coherency constrained models on the Aff-wild dataset. Temporal coherency constrained models outperformed the unconstrained models by a significant margin. This performance improvement was also observed in the case of cross-dataset evaluation results on AFEW -VA database. Most notably, temporally coherent visual features produced state-of-the-art performance on the Aff-wild benchmark without using additional inputs such as facial landmarks.