Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection

Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection
复制标题

DOI:
10.1109/tpami.2022.3180392
复制
发表时间:
2023-03-01
影响因子:
23.6
通讯作者:
Yang, Jian
Yang, Jian
中科院分区:
计算机科学1区
文献类型:
--
作者:
Li, Xiang;Lv, Chengqi;Yang, Jian

文献摘要

被引文献

相似文献

目标检测是一项基本的计算机视觉任务,它同时预测感兴趣目标的类别和位置。近年来,单级(也称为“密集”)检测器由于其简单的流水线和友好的终端设备应用而比两级检测器受到更多的关注。密集对象检测器基本上将对象检测公式化为密集分类和定位(即,边界框回归)。分类通常通过Focal Loss进行优化,并且框位置通常在Dirac delta分布下学习。密集检测器的一个新趋势是引入一个单独的预测分支来估计定位的质量,这有助于分类以提高检测性能。本文深入研究了上述三个基本要素的表示:质量估计,分类和定位。在现有的实践中发现了三个问题,包括(1)在训练和推理之间不一致地使用质量估计和分类,(2)用于定位的不灵活的Dirac delta分布,以及(3)用于精确质量估计的缺乏和隐含的指导。为了解决这些问题,我们为这些元素设计了新的表示。具体来说,我们合并到类预测向量的质量估计,以形成一个联合表示,使用一个向量来表示任意分布的框位置,并提取判别特征描述符的分布向量更可靠的质量估计。改进后的表示方法消除了不一致性风险,准确刻画了真实的数据中的柔性分布,但包含连续的标签,超出了焦点损失的范围。然后,我们提出了广义焦点损失(GFocal),将焦点损失从离散形式推广到连续形式,以实现成功的优化。大量的实验表明,我们的方法的有效性,而不牺牲效率的训练和推理。基于GFocal,我们在移动的设置下构建了一个相当快速和轻量级的检测器NanoDet,它比缩放的YoloV 4-Tiny高1.8 AP,快2倍,小6倍。
Object detection is a fundamental computer vision task that simultaneously predicts the category and localization of the targets of interest. Recently one-stage (also termed "dense") detectors have gained much attention over two-stage ones due to their simple pipeline and friendly application to end devices. Dense object detectors basically formulate object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for dense detectors is to introduce an individual prediction branch to estimate the quality of localization, which facilitates the classification to improve detection performance. This paper delves into the representations of the above three fundamental elements: quality estimation, classification and localization. Three problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, (2) the inflexible Dirac delta distribution for localization, and (3) the deficient and implicit guidance for accurate quality estimation. To address these problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, use a vector to represent arbitrary distribution of box locations, and extract discriminant feature descriptors from the distribution vector for more reliable quality estimation. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain continuous labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFocal) that generalizes Focal Loss from its discrete form to the continuous version for successful optimization. Extensive experiments demonstrate the effectiveness of our method, without sacrificing the efficiency both in training and inference. Based on GFocal, we construct a considerably fast and lightweight detector termed NanoDet under mobile settings, which is 1.8 AP higher, 2x faster and 6x smaller than scaled YoloV4-Tiny.