Cascade R-CNN: High Quality Object Detection and Instance Segmentation

Cascade R-CNN: High Quality Object Detection and Instance Segmentation
复制标题

DOI:
10.1109/tpami.2019.2956516
复制
发表时间:
2021-05-01
影响因子:
23.6
通讯作者:
Vasconcelos, Nuno
Vasconcelos, Nuno
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cai, Zhaowei;Vasconcelos, Nuno

文献摘要

被引文献

相似文献

在目标检测中,交集与并集(IOU)阈值经常被用来定义正/负。用于训练检测器的阈值定义了它的质量。虽然通常使用的阈值为0.5会导致噪声(低质量)检测,但阈值越大,检测性能就越差。这种高质量检测的悖论有两个原因:1)由于大阈值的正样本消失而导致的过度拟合,以及2)检测器和测试假设之间的推理时间质量不匹配。为了解决这些问题,提出了一种多级目标检测体系结构Cascade R-CNN,该体系结构由一系列检测器组成,并通过不断增加的IOU阈值进行训练。使用检测器的输出作为下一步的训练集,对检测器进行顺序训练。这种重采样逐渐提高了假设质量,保证了所有检测器的相同大小的正训练集,并最大限度地减少了过拟合。同样的级联应用于推理,以消除假设和检测器之间的质量不匹配。没有铃声或口哨的级联R-CNN的实施实现了对COCO数据集的最先进的性能,并显著提高了对通用和特定对象数据集的高质量检测,包括VOC、Kitti、CityPerson和WiderFace。最后,将级联R-CNN推广到实例分割中,并对MASK R-CNN进行了改进。
In object detection, the intersection over union (IoU) threshold is frequently used to define positives/negatives. The threshold used to train a detector defines its quality. While the commonly used threshold of 0.5 leads to noisy (low-quality) detections, detection performance frequently degrades for larger thresholds. This paradox of high-quality detection has two causes: 1) overfitting, due to vanishing positive samples for large thresholds, and 2) inference-time quality mismatch between detector and test hypotheses. A multi-stage object detection architecture, the Cascade R-CNN, composed of a sequence of detectors trained with increasing IoU thresholds, is proposed to address these problems. The detectors are trained sequentially, using the output of a detector as training set for the next. This resampling progressively improves hypotheses quality, guaranteeing a positive training set of equivalent size for all detectors and minimizing overfitting. The same cascade is applied at inference, to eliminate quality mismatches between hypotheses and detectors. An implementation of the Cascade R-CNN without bells or whistles achieves state-of-the-art performance on the COCO dataset, and significantly improves high-quality detection on generic and specific object datasets, including VOC, KITTI, CityPerson, and WiderFace. Finally, the Cascade R-CNN is generalized to instance segmentation, with nontrivial improvements over the Mask R-CNN.