Temporally Coherent Visual Representations for Dimensional Affect Recognition

Temporally Coherent Visual Representations for Dimensional Affect Recognition
复制标题

DOI:
10.1109/acii.2019.8925529
复制
发表时间:
2019-09
期刊:
2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII)
影响因子:
--
通讯作者:
M. Tellamekala;M. Valstar
M. Tellamekala;M. Valstar
中科院分区:
其他
文献类型:
--
作者:
M. Tellamekala;M. Valstar

文献摘要

被引文献

相似文献

近年来,端到端监督表示学习和回归的成功已经将连续情感识别方法的主要焦点转移到直接从大规模标记数据集学习视觉表示。监督表示学习对于具有明确标签的学习任务非常有效。但是标注维度情感数据本质上是主观的,这往往会导致模糊的标签。仅仅依赖于这种模糊的标签进行表示学习不会产生具有良好泛化能力的鲁棒特征,正如我们将在这项工作中展示的那样。为了解决这个根本问题,我们建议应用一个约束表示学习方法,鼓励潜在的功能是不太敏感的模糊的情感注释,通过使用一个通用的表示学习之前称为“时间一致性”或“时间平滑度”。这种方法通过向监督学习损失函数添加一阶时间相干性正则化约束来迫使潜在特征在时间上是相干的。为了评估时间相干表示的效用,我们在Aff-wild数据集上训练了无约束和时间相干约束模型。时间一致性约束模型优于无约束模型的显着利润率。在AFEW-VA数据库的交叉数据集评价结果中也观察到这种性能改善。最值得注意的是,时间上连贯的视觉特征在Aff-wild基准测试中产生了最先进的性能,而无需使用额外的输入,如面部标志。
The success of end-to-end supervised representation learning and regression in recent years has shifted the main focus of continuous emotion recognition approaches to learning visual representations directly from large scale labeled datasets. Supervised representation learning is highly effective for learning tasks with unambiguous labels. But annotating dimensional affect data is inherently subjective which more often than not leads to ambiguous labels. Relying only on such ambiguous labels for representation learning does not result in robust features with good generalization capacity, as we will show in this work. To address this fundamental problem, we propose to apply a constrained representation learning method that encourages the latent features to be less sensitive to ambiguous emotion annotations by using a generic representation learning prior called ‘temporal coherency’ or ‘temporal smoothness’. This approach forces the latent features to be temporally coherent, by adding a first-order temporal coherency regularization constraint to the supervised learning loss function. To assess the utility of temporally coherent representations, we trained both unconstrained and temporal coherency constrained models on the Aff-wild dataset. Temporal coherency constrained models outperformed the unconstrained models by a significant margin. This performance improvement was also observed in the case of cross-dataset evaluation results on AFEW -VA database. Most notably, temporally coherent visual features produced state-of-the-art performance on the Aff-wild benchmark without using additional inputs such as facial landmarks.