Modelling Stochastic Context of Audio-Visual Expressive Behaviour With Affective Processes

Modelling Stochastic Context of Audio-Visual Expressive Behaviour With Affective Processes
复制标题

DOI:
10.1109/taffc.2022.3157141
复制
发表时间:
2023-07
影响因子:
11.2
通讯作者:
M. Tellamekala;T. Giesbrecht;M. Valstar
M. Tellamekala;T. Giesbrecht;M. Valstar
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Tellamekala;T. Giesbrecht;M. Valstar

文献摘要

相似文献

在自然条件下从视听信号中识别明显的情感仍然是一个悬而未决的问题。基于循环模型的现有方法,或者使用自注意力在特征级别对上下文依赖关系进行建模的方法,无法对在不同抽象级别微妙发生的长期依赖关系进行建模。情感过程已经成为一种通过概率性全局潜在变量来建模时间动态的新范式,该变量捕获上下文并在输出中引入依赖性,显示出卓越的性能且复杂性极低。尽管在视觉数据上取得了令人印象深刻的结果,但情感过程在音频数据领域仍未得到探索,众所周知,情感过程对情绪的感知有至关重要的影响。在本文中,我们首先重新审视情感过程并将其扩展到语音领域,确定其有效训练的关键组成部分和学习程序。然后,我们使用特定于模态的上下文编码器将情感过程扩展到视听情感识别。最后,我们提出了情感过程在协作机器学习领域的一种新颖应用,用于使用稀疏的人类监督在视频中传播情感标签。我们进行了广泛的消融研究,确定了情感过程成功背后的主要组成部分,并与各种数据集中的现有作品进行了比较。
Recognising apparent emotion from audio-visual signals in naturalistic conditions remains an open problem. Existing methods that build on recurrent models, or in the modelling of contextual dependencies at the feature level using self-attention fail to model the long-term dependencies that subtly occur at different levels of abstraction. Affective Processes have emerged as a novel paradigm to the modelling of temporal dynamics through a probabilistic global latent variable that captures context and induces dependencies in the outputs, showing superior performance with little complexity. Despite its impressive results on visual data, Affective Processes remain unexplored in the domain of audio data, known to crucially influence the perception of emotions. In this paper, we first revisit and extend Affective Processes to the speech domain, identifying the key components and learning procedures for their efficient training. We then extend Affective Processes to audio-visual affect recognition, using modality-specific context encoders. Finally, we propose a novel application of Affective Processes in the domain of Cooperative Machine Learning for propagating affect labels in videos using sparse human supervision. We conduct extensive ablation studies, identifying the main components behind the success of Affective Processes, as well as comparisons against existing works in a variety of datasets.