Learning Facial Expressions with 3D Mesh Convolutional Neural Network

Learning Facial Expressions with 3D Mesh Convolutional Neural Network
复制标题

DOI:
10.1145/3200572
复制
发表时间:
2018-11
期刊:
ACM Transactions on Intelligent Systems and Technology (TIST)
影响因子:
--
通讯作者:
Hai Jin;Yuanfeng Lian;Jing Hua
Hai Jin;Yuanfeng Lian;Jing Hua
中科院分区:
其他
文献类型:
--
作者:
Hai Jin;Yuanfeng Lian;Jing Hua

文献摘要

相似文献

让机器理解人类的表达可以在人机交互中实现各种有用的应用。在本文中,我们提出了一种采用 3D 网格卷积神经网络 (3DMCNN) 的新型面部表情识别方法以及视觉分析引导的 3DMCNN 设计和优化方案。我们首先通过 RGBD 相机重建具有面部表情的主体的 3D 面部模型,然后计算表面的几何属性。我们没有使用常规的卷积神经网络 (CNN) 来学习面部图像的强度,而是使用 3DMCNN 对 3D 模型表面的几何属性进行卷积。我们设计了一种基于测地距离的卷积方法来克服人脸表面网格不规则采样带来的困难。我们进一步提出交互式视觉分析,目的是设计和修改网络,以分析学习到的特征并聚类 3DMCNN 中的相似节点。通过去除网络中低活跃度的节点,网络的性能得到极大的提高。我们通过交互式可视化网络的每一层,将我们的方法与常规的基于 CNN 的方法进行比较,并通过研究代表性案例来分析我们方法的有效性。在公共数据集上进行测试,我们的方法比传统的基于图像的 CNN 和其他 3D CNN 实现了更高的识别精度。所提出的框架,包括 3DMCNN 和 CNN 的交互式视觉分析,可以扩展到其他应用程序。
Making machines understand human expressions enables various useful applications in human-machine interaction. In this article, we present a novel facial expression recognition approach with 3D Mesh Convolutional Neural Networks (3DMCNN) and a visual analytics-guided 3DMCNN design and optimization scheme. From an RGBD camera, we first reconstruct a 3D face model of a subject with facial expressions and then compute the geometric properties of the surface. Instead of using regular Convolutional Neural Networks (CNNs) to learn intensities of the facial images, we convolve the geometric properties on the surface of the 3D model using 3DMCNN. We design a geodesic distance-based convolution method to overcome the difficulties raised from the irregular sampling of the face surface mesh. We further present interactive visual analytics for the purpose of designing and modifying the networks to analyze the learned features and cluster similar nodes in 3DMCNN. By removing low-activity nodes in the network, the performance of the network is greatly improved. We compare our method with the regular CNN-based method by interactively visualizing each layer of the networks and analyze the effectiveness of our method by studying representative cases. Testing on public datasets, our method achieves a higher recognition accuracy than traditional image-based CNN and other 3D CNNs. The proposed framework, including 3DMCNN and interactive visual analytics of the CNN, can be extended to other applications.