Towards Long-Tailed 3D Detection

Towards Long-Tailed 3D Detection
复制标题

DOI:
10.48550/arxiv.2211.08691
复制
发表时间:
2022-11
期刊:
--
影响因子:
--
通讯作者:
Neehar Peri;Achal Dave;Deva Ramanan;Shu Kong
Neehar Peri;Achal Dave;Deva Ramanan;Shu Kong
中科院分区:
其他
文献类型:
--
作者:
Neehar Peri;Achal Dave;Deva Ramanan;Shu Kong

文献摘要

相似文献

当代自动驾驶汽车 (AV) 基准测试拥有用于训练 3D 探测器的先进技术,特别是在大规模激光雷达数据上。令人惊讶的是,尽管语义类别标签自然遵循长尾分布,但当代基准仅关注少数常见类别(例如行人和汽车),而忽略了尾部的许多罕见类别(例如碎片和婴儿车)。然而,AV 仍然必须检测稀有类别以确保安全运行。此外,语义类通常在层次结构中组织,例如,诸如儿童和建筑工人之类的尾部类可以说是行人的子类。然而,这种层次关系经常被忽视,这可能会导致对性能的误导性估计并错失算法创新的机会。我们通过正式研究长尾 3D 检测 (LT3D) 问题来解决这些挑战,该问题对所有类别(包括尾部类别)进行评估。我们对流行的 3D 检测代码库(例如 CenterPoint 和 PointPillars)进行评估和创新,使它们适用于 LT3D。我们开发了层次损失,促进常见类与稀有类之间的特征共享,并改进了检测指标,将部分信用奖励给尊重层次结构的“合理”错误(例如,将儿童误认为成人)。最后,我们指出,通过 RGB 图像与 LiDAR 的多模态融合,细粒度尾部类别精度得到了特别提高;简而言之,仅从稀疏(激光雷达)几何形状中识别小型细粒度类别具有挑战性,这表明多模态线索对于长尾 3D 检测至关重要。我们的修改将所有类别的 AP 平均准确率提高了 5%,并显着提高了稀有类别的 AP(例如,婴儿车 AP 从 3.6 提高到 31.6)!我们的代码位于 https://github.com/neeharperi/LT3D
Contemporary autonomous vehicle (AV) benchmarks have advanced techniques for training 3D detectors, particularly on large-scale lidar data. Surprisingly, although semantic class labels naturally follow a long-tailed distribution, contemporary benchmarks focus on only a few common classes (e.g., pedestrian and car) and neglect many rare classes in-the-tail (e.g., debris and stroller). However, AVs must still detect rare classes to ensure safe operation. Moreover, semantic classes are often organized within a hierarchy, e.g., tail classes such as child and construction-worker are arguably subclasses of pedestrian. However, such hierarchical relationships are often ignored, which may lead to misleading estimates of performance and missed opportunities for algorithmic innovation. We address these challenges by formally studying the problem of Long-Tailed 3D Detection (LT3D), which evaluates on all classes, including those in-the-tail. We evaluate and innovate upon popular 3D detection codebases, such as CenterPoint and PointPillars, adapting them for LT3D. We develop hierarchical losses that promote feature sharing across common-vs-rare classes, as well as improved detection metrics that award partial credit to"reasonable"mistakes respecting the hierarchy (e.g., mistaking a child for an adult). Finally, we point out that fine-grained tail class accuracy is particularly improved via multimodal fusion of RGB images with LiDAR; simply put, small fine-grained classes are challenging to identify from sparse (lidar) geometry alone, suggesting that multimodal cues are crucial to long-tailed 3D detection. Our modifications improve accuracy by 5% AP on average for all classes, and dramatically improve AP for rare classes (e.g., stroller AP improves from 3.6 to 31.6)! Our code is available at https://github.com/neeharperi/LT3D