Keep an eye on faces: Robust face detection with heatmap-Assisted spatial attention and scale-Aware layer attention

Keep an eye on faces: Robust face detection with heatmap-Assisted spatial attention and scale-Aware layer attention
复制标题

DOI:
10.1016/j.patcog.2023.109553
复制
发表时间:
2023-08
期刊:
Pattern Recognit.
影响因子:
--
通讯作者:
Lei Ju;J. Kittler;M. A. Rana;Wankou Yang;Zhenhua Feng
Lei Ju;J. Kittler;M. A. Rana;Wankou Yang;Zhenhua Feng
中科院分区:
其他
文献类型:
--
作者:
Lei Ju;J. Kittler;M. A. Rana;Wankou Yang;Zhenhua Feng

文献摘要

相似文献

现代基于锚点的人脸检测器使用大容量网络和广泛的锚点设置来学习判别特征。尽管结果令人鼓舞,但也并非没有问题。首先,大多数锚点从背景中提取冗余特征。因此,性能的提高是以不成比例的计算复杂性为代价的。其次,预测的人脸框仅由由预定义的正、负和忽略锚点监督的分类器来区分。该策略可能会忽略在推理过程中被标记为负/被忽略的锚群体的潜在贡献,仅仅是因为它们的初始化较差,尽管它们可以很好地回归到目标。换句话说,真正的阳性和代表性特征可能会被不可靠的置信度分数过滤掉。为了解决第一个问题并实现更有效的人脸检测,我们提出了热图辅助空间注意力(HSA)模块和尺度感知层注意力(SLA)模块,以使用较低的计算成本提取信息特征。具体来说,SLA 融合了所有特征金字塔层的信息,并自适应加权以去除冗余层。 HSA 预测重塑的高斯热图,并利用它通过更好地突出面部区域来促进空间特征选择。为了更可靠的决策,我们通过投票合并预测的热图分数和分类结果。由于我们的热图分数是基于到面部中心的距离,因此它们能够保留所有回归良好的锚点。在几个著名基准上获得的实验证明了所提出方法的优点。
Modern anchor-based face detectors learn discriminative features using large-capacity networks and extensive anchor settings. In spite of their promising results, they are not without problems. First, most anchors extract redundant features from the background. As a consequence, the performance improvements are achieved at the expense of a disproportionate computational complexity. Second, the predicted face boxes are only distinguished by a classifier supervised by pre-defined positive, negative and ignored anchors. This strategy may ignore potential contributions from cohorts of anchors labeled negative/ignored during inference simply because of their inferior initialisation, although they can regress well to a target. In other words, true positives and representative features may get filtered out by unreliable confidence scores. To deal with the first concern and achieve more efficient face detection, we propose a Heatmap-assisted Spatial Attention (HSA) module and a Scale-aware Layer Attention (SLA) module to extract informative features using lower computational costs. To be specific, SLA incorporates the information from all the feature pyramid layers, weighted adaptively to remove redundant layers. HSA predicts a reshaped Gaussian heatmap and employs it to facilitate a spatial feature selection by better highlighting facial areas. For more reliable decision-making, we merge the predicted heatmap scores and classification results by voting. Since our heatmap scores are based on the distance to the face centres, they are able to retain all the well-regressed anchors. The experiments obtained on several well-known benchmarks demonstrate the merits of the proposed method.