Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection

Exploiting Semantic Embedding and Visual Feature for Facial Action Unit Detection
复制标题

DOI:
10.1109/cvpr46437.2021.01034
复制
发表时间:
2021-06
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Huiyuan Yang;L. Yin;Yi Zhou;Jiuxiang Gu
Huiyuan Yang;L. Yin;Yi Zhou;Jiuxiang Gu
中科院分区:
其他
文献类型:
--
作者:
Huiyuan Yang;L. Yin;Yi Zhou;Jiuxiang Gu

文献摘要

被引文献

相似文献

最近关于检测面部动作单元(AU)的研究利用了辅助信息(即面部标志、AU和表情之间的关系、网络面部图像等),以提高AU检测性能。截至目前,尚未针对此类任务探索 AU 的语义信息。事实上,AU 语义描述比单独的二进制 AU 标签提供了更多的信息,因此我们建议利用语义嵌入和视觉特征(SEV-Net)进行 AU 检测。更具体地说,AU语义嵌入是通过AU内和AU间注意模块获得的,其中AU内注意模块捕获描述单个AU的每个句子中单词之间的关系,而AU间注意模块则关注这些句子之间的关系。然后将学习到的 AU 语义嵌入用作通过跨模态注意力网络生成注意力图的指导。生成的跨模态注意力图进一步用作聚合特征的权重。我们提出的方法是独特的,因为语义特征被利用为此类方法中的第一个。该方法已在三个公共 AU 编码的面部表情数据库上进行了评估,并取得了比最先进的同行方法更优越的性能。
Recent study on detecting facial action units (AU) has utilized auxiliary information (i.e., facial landmarks, relationship among AUs and expressions, web facial images, etc.), in order to improve the AU detection performance. As of now, no semantic information of AUs has yet been explored for such a task. As a matter of fact, AU semantic descriptions provide much more information than the binary AU labels alone, thus we propose to exploit the Semantic Embedding and Visual feature (SEV-Net) for AU detection. More specifically, AU semantic embeddings are obtained through both Intra-AU and Inter-AU attention modules, where the Intra-AU attention module captures the relation among words within each sentence that describes individual AU, and the Inter-AU attention module focuses on the relation among those sentences. The learned AU semantic embeddings are then used as guidance for the generation of attention maps through a cross-modality attention network. The generated cross-modality attention maps are further used as weights for the aggregated feature. Our proposed method is unique in that the semantic features are exploited as the first of this kind. The approach has been evaluated on three public AU-coded facial expression databases, and has achieved a superior performance than the state-of-the-art peer methods.