Context prior-based with residual learning for face detection: A deep convolutional encoder-decoder network

Context prior-based with residual learning for face detection: A deep convolutional encoder-decoder network
复制标题

基于上下文先验的人脸检测残差学习:深度卷积编码器-解码器网络

DOI:
10.1016/j.image.2020.115948
复制
发表时间:
2020-10-01
影响因子:
3.5
通讯作者:
Chen, Ziyu
Chen, Ziyu
中科院分区:
工程技术2区
文献类型:
--
作者:
Zhou, Zexun;He, Zhongshi;Chen, Ziyu

文献摘要

被引文献

相似文献

在安防领域,室外监控摄像机拍摄的图像中,人脸通常具有模糊、遮挡、姿态多样、体积小等特点,受摄像机姿态和距离、天气条件等外部环境的影响,可以说是自然图像中人脸检测的难题。为了解决这个问题,我们提出了一种名为特征层次编码器-解码器网络(FHEDN)的深度卷积神经网络。它的动机是从上下文语义信息和多尺度人脸检测机制的两个观察。该网络是一种变尺度结构的单级网络,由编码子网络和解码子网络组成。基于人脸周围的上下文语义信息是辅助人脸检测的假设,我们引入了一种残差机制,将上下文先验信息融合到人脸特征中,并制定学习链来训练每个编码器-解码器对。此外,我们还讨论了实现细节中的一些重要因素,如训练数据集的分布、特征层次的规模、锚盒的大小等,它们对最终网络的检测性能有一定的影响。与一些国家的最先进的算法相比,我们的方法取得了可喜的性能在流行的基准测试,包括AFW,PASCAL FACE,FDDB,和WIDER FACE。因此,所提出的方法可以有效地实现和常规应用于检测人脸严重遮挡和任意姿态变化的无约束场景。我们的代码和结果可以在https://github.com/zzxcoder/EvaluationFHEDN上找到。
In the field of security, faces are usually blurry, occluded, diverse pose and small in the image captured by an outdoor surveillance camera, which is affected by the external environment such as the camera pose and range, weather conditions, etc. It can be described as a problem of hard face detection in natural images. To solve this problem, we propose a deep convolutional neural network named feature hierarchy encoder-decoder network (FHEDN). It is motivated by two observations from contextual semantic information and the mechanism of multi-scale face detection. The proposed network is a scale-variant style architecture and single stage, which are composed of encoder and decoder subnetworks. Based on the assumption that contextual semantic information around face being auxiliary to detect faces, we introduce a residual mechanism to fuse context prior-based information into face feature and formulate the learning chain to train each encoder-decoder pair. In addition, we discuss some important factors in implement details such as the distribution of training dataset, the scale of feature hierarchy, and anchor box size, etc. They have some impact on the detection performance of the final network. Compared with some state-of-the-art algorithms, our method achieves promising performance on the popular benchmarks including AFW, PASCAL FACE, FDDB, and WIDER FACE. Consequently, the proposed approach can be efficiently implemented and routinely applied to detect faces with severe occlusion and arbitrary pose variations in unconstrained scenes. Our code and results are available on https://github.com/zzxcoder/EvaluationFHEDN.