Adaptive Multimodal Fusion for Facial Action Units Recognition

Adaptive Multimodal Fusion for Facial Action Units Recognition
复制标题

DOI:
10.1145/3394171.3413538
复制
发表时间:
2020-10
期刊:
Proceedings of the 28th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
Huiyuan Yang;Taoyue Wang;L. Yin
Huiyuan Yang;Taoyue Wang;L. Yin
中科院分区:
其他
文献类型:
--
作者:
Huiyuan Yang;Taoyue Wang;L. Yin

文献摘要

相似文献

多模态面部动作单元(Au)识别旨在构建能够处理、关联和整合来自多个模态(即,来自视觉传感器的2D图像、来自3D成像的3D几何形状以及来自红外传感器的热图像)。虽然多模型数据可以提供丰富的信息,但在从多模态数据中学习时必须解决两个挑战:1)模型必须捕获复杂的跨模态交互,以便有效地利用附加和互信息; 2)模型必须在测试期间意外的数据损坏的情况下足够鲁棒,如果某个模态丢失或有噪声。在本文中,我们提出了一种新的适应性多模态F-检验方法(AMF)的Au检测,学习选择最相关的特征表示从不同的模态的重采样过程的条件下的特征评分模块。特征评分模块被设计为允许评估从多个模态学习的特征的质量。因此,AMF能够自适应地选择更具鉴别力的特征,从而增加对丢失或损坏模态的鲁棒性。此外,为了缓解过拟合问题,使模型更好地推广测试数据,剪切开关多模态数据增强方法的设计,通过该方法,一个随机块被剪切和切换跨多个模态。我们对两个公开的多模态Au数据集BP 4D和BP 4D+进行了深入的研究,结果证明了所提出的方法的有效性。各种情况下的消融研究也表明,我们的方法仍然强大的测试过程中丢失或嘈杂的方式。
Multimodal facial action units (AU) recognition aims to build models that are capable of processing, correlating, and integrating information from multiple modalities (i.e., 2D images from a visual sensor, 3D geometry from 3D imaging, and thermal images from an infrared sensor). Although the multimodel data can provide rich information, there are two challenges that have to be addressed when learning from multimodal data: 1) the model must capture the complex cross-modal interactions in order to utilize the additional and mutual information effectively; 2) the model must be robust enough in the circumstance of unexpected data corruptions during testing, in case of a certain modality missing or being noisy. In this paper, we propose a novel A daptive M ultimodal F usion method (AMF ) for AU detection, which learns to select the most relevant feature representations from different modalities by a re-sampling procedure conditioned on a feature scoring module. The feature scoring module is designed to allow for evaluating the quality of features learned from multiple modalities. As a result, AMF is able to adaptively select more discriminative features, thus increasing the robustness to missing or corrupted modalities. In addition, to alleviate the over-fitting problem and make the model generalize better on the testing data, a cut-switch multimodal data augmentation method is designed, by which a random block is cut and switched across multiple modalities. We have conducted a thorough investigation on two public multimodal AU datasets, BP4D and BP4D+, and the results demonstrate the effectiveness of the proposed method. Ablation studies on various circumstances also show that our method remains robust to missing or noisy modalities during tests.