Improving Face Recognition from Caption Supervision with Multi-Granular Contextual Feature Aggregation

Improving Face Recognition from Caption Supervision with Multi-Granular Contextual Feature Aggregation
复制标题

DOI:
10.1109/ijcb57857.2023.10448749
复制
发表时间:
2023-08
期刊:
2023 IEEE International Joint Conference on Biometrics (IJCB)
影响因子:
--
通讯作者:
M. Hasan;N. Nasrabadi
M. Hasan;N. Nasrabadi
中科院分区:
其他
文献类型:
--
作者:
M. Hasan;N. Nasrabadi

文献摘要

相似文献

我们引入字幕引导人脸识别(CGFR)作为一种新框架,以提高商用现成(COTS)人脸识别(FR)系统的性能。与将软生物特征识别(例如面部标记、性别和年龄)与面部图像相结合相比,在这项工作中,我们使用面部检查者提供的面部描述作为辅助信息。然而,由于模态的异构性,通过直接融合文本和面部特征来提高性能非常具有挑战性,因为两者位于不同的嵌入空间。在本文中,我们提出了一种上下文特征聚合模块(CFAM),它通过有效利用细粒度的单词区域交互和全局图像标题关联来解决这个问题。具体来说,CFAM采用自注意力和交叉注意力方案来分别改善图像和文本特征之间的模态内和模态间关系。此外,我们设计了一个文本特征细化模块(TFRM),通过更新上下文嵌入来细化预训练的 BERT 编码器的文本特征。该模块通过跨模态投影损失增强了文本特征的判别能力,并通过合并视觉语义对齐损失将单词和标题嵌入与视觉特征重新对齐。我们在两个人脸识别模型(ArcFace 和 AdaFace)上实现了所提出的 CGFR 框架,并评估了其在多模态 CelebA-HQ 数据集上的性能。我们的框架显着提高了 ArcFace 在 1:1 验证和 1:N 识别协议中的性能。
We introduce caption-guided face recognition (CGFR) as a new framework to improve the performance of commercial-off-the-shelf (COTS) face recognition (FR) systems. In contrast to combining soft biometrics (e.g., facial marks, gender, and age) with face images, in this work, we use facial descriptions provided by face examiners as a piece of auxiliary information. However, due to the heterogeneity of the modalities, improving the performance by directly fusing the textual and facial features is very challenging, as both lie in different embedding spaces. In this paper, we propose a contextual feature aggregation module (CFAM) that addresses this issue by effectively exploiting the fine-grained word-region interaction and global image-caption association. Specifically, CFAM adopts a self-attention and a cross-attention scheme for improving the intra-modality and inter-modality relationship between the image and textual features, respectively. Additionally, we design a textual feature refinement module (TFRM) that refines the textual features of the pre-trained BERT encoder by updating the contextual embeddings. This module enhances the discriminative power of textual features with a cross-modal projection loss and realigns the word and caption embeddings with visual features by incorporating a visual-semantic alignment loss. We implemented the proposed CGFR framework on two face recognition models (ArcFace and AdaFace) and evaluated its performance on the Multi-Modal CelebA-HQ dataset. Our framework significantly improves the performance of ArcFace in both 1:1 verification and 1:N identification protocol.