Multimodal Framework for Analyzing the Affect of a Group of People

Multimodal Framework for Analyzing the Affect of a Group of People
复制标题

DOI:
10.1109/tmm.2018.2818015
复制
发表时间:
2018-03
影响因子:
7.3
通讯作者:
Xiaohua Huang;Abhinav Dhall;R. Goecke;M. Pietikäinen;Guoying Zhao
Xiaohua Huang;Abhinav Dhall;R. Goecke;M. Pietikäinen;Guoying Zhao
中科院分区:
计算机科学1区
文献类型:
--
作者:
Xiaohua Huang;Abhinav Dhall;R. Goecke;M. Pietikäinen;Guoying Zhao

文献摘要

被引文献

相似文献

随着多媒体和万维网的进步,用户在互联网上的社交网络平台上上传数百万张图像和视频。从自动人类行为理解的角度来看,分析和建模参与这些图像中的社会事件的人群所表现出的影响是有意义的。然而,由于室内和室外环境的变化,对多人表达的影响进行分析是具有挑战性的。最近,一些有趣的工作已经研究了基于人脸的群体级情绪识别(格尔)。在本文中,我们提出了一个多模态框架,以提高情感分析能力的格尔在具有挑战性的环境。具体来说,对于编码一个人的信息在一个组级的图像,我们首先提出了一个信息聚合方法,用于生成面部,上身和场景的特征描述。稍后,我们将重新审视本地化的多核学习,以融合面部,上身和场景信息,以应对具有挑战性的环境。在两个具有挑战性的组级情感数据库(HAPPEI和GAFF)上进行了深入的实验,以研究面部,上身,场景信息和多模态框架的作用。实验结果表明,多模态框架实现了良好的性能的GER。
With the advances in multimedia and the world wide web, users upload millions of images and videos everyone on social networking platforms on the Internet. From the perspective of automatic human behavior understanding, it is of interest to analyze and model the affects that are exhibited by groups of people who are participating in social events in these images. However, the analysis of the affect that is expressed by multiple people is challenging due to the varied indoor and outdoor settings. Recently, a few interesting works have investigated face-based group-level emotion recognition (GER). In this paper, we propose a multimodal framework for enhancing the affective analysis ability of GER in challenging environments. Specifically, for encoding a person's information in a group-level image, we first propose an information aggregation method for generating feature descriptions of face, upper body, and scene. Later, we revisit localized multiple kernel learning for fusing face, upper body, and scene information for GER against challenging environments. Intensive experiments are performed on two challenging group-level emotion databases (HAPPEI and GAFF) to investigate the roles of the face, upper body, scene information, and the multimodal framework. Experimental results demonstrate that the multimodal framework achieves promising performance for GER.