Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation

Low-Resolution Face Recognition in the Wild via Selective Knowledge Distillation
复制标题

通过选择性知识蒸馏进行野外低分辨率人脸识别

DOI:
10.1109/tip.2018.2883743
复制
发表时间:
2019-04-01
影响因子:
10.6
通讯作者:
Li, Jia
Li, Jia
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ge, Shiming;Zhao, Shengwei;Li, Jia

文献摘要

被引文献

相似文献

通常,在野外部署人脸识别模型需要以极低的计算成本识别低分辨率的人脸。为了解决这个问题,一个可行的解决方案是压缩复杂的人脸模型,以最小的性能下降为代价实现更高的速度和更低的内存。受此启发,本文提出了一种基于选择性知识蒸馏的低分辨率人脸识别学习方法。在这种方法中,首先初始化两流卷积神经网络(CNN),分别识别具有教师流和学生流的高分辨率人脸和分辨率下降的人脸。教师流用一个复杂的CNN表示,以实现高精度的识别;学生流用一个简单得多的CNN表示,以实现低复杂度的识别。为了避免学生流的显著性能下降,我们通过解决稀疏图优化问题,有选择地从教师流中提取最具信息量的面部特征,然后将其用于正则化学生流的微调过程。通过这种方式,学生流实际上是通过在有限的计算资源下同时处理两项任务来训练的:通过特征回归近似最具信息量的面部线索,以及通过低分辨率面部分类恢复缺失的面部线索。实验结果表明,学生流在识别低分辨率人脸方面表现出色,仅占用0.15 mb内存,在CPU上每秒运行418张人脸,在GPU上每秒运行9433张人脸。
Typically, the deployment of face recognition models in the wild needs to identify low-resolution faces with extremely low computational cost. To address this problem, a feasible solution is compressing a complex face model to achieve higher speed and lower memory at the cost of minimal performance drop. Inspired by that, this paper proposes a learning approach to recognize low-resolution faces via selective knowledge distillation. In this approach, a two-stream convolutional neural network (CNN) is first initialized to recognize high-resolution faces and resolution-degraded faces with a teacher stream and a student stream, respectively. The teacher stream is represented by a complex CNN for high-accuracy recognition, and the student stream is represented by a much simpler CNN for low-complexity recognition. To avoid significant performance drop at the student stream, we then selectively distil the most informative facial features from the teacher stream by solving a sparse graph optimization problem, which are then used to regularize the fine-tuning process of the student stream. In this way, the student stream is actually trained by simultaneously handling two tasks with limited computational resources: approximating the most informative facial cues via feature regression, and recovering the missing facial cues via low-resolution face classification. Experimental results show that the student stream performs impressively in recognizing low-resolution faces and costs only 0.15-MB memory and runs at 418 faces per second on CPU and 9433 faces per second on GPU.