WALT: Watch And Learn 2D amodal representation from Time-lapse imagery

WALT: Watch And Learn 2D amodal representation from Time-lapse imagery
复制标题

DOI:
10.1109/cvpr52688.2022.00914
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Dinesh Reddy Narapureddy;R. Tamburo;S. Narasimhan
Dinesh Reddy Narapureddy;R. Tamburo;S. Narasimhan
中科院分区:
其他
文献类型:
--
作者:
Dinesh Reddy Narapureddy;R. Tamburo;S. Narasimhan

文献摘要

相似文献

当前用于对象检测、分割和跟踪的方法在存在严重遮挡的忙碌城市环境中失败。标记的真实的遮挡数据是稀缺的(即使在大型数据集中),合成数据留下了一个域的差距,使得很难明确建模和学习遮挡。在这项工作中,我们提出了最好的真实的和合成世界的自动遮挡监督使用一个大的现成的数据源:延时图像从固定的网络摄像头观察街道交叉口数周,数月,甚至数年。我们引入了一个新的数据集,观看和学习延时(WALT),由12个(4K和1080 p)摄像头组成,在一年内捕捉城市环境。我们利用这个真实的数据,以一种新颖的方式自动挖掘一个大的未被遮挡的对象,然后将它们组合在相同的视图中生成遮挡。这种纵向自我监督足够强大,可以让非模态网络学习对象-遮挡物-遮挡层表示。我们将展示如何加快未被遮挡的对象的发现,并在此发现的信心,训练遮挡对象的速度和准确性。经过几天的观察和自动学习,这种方法在检测和分割被遮挡的人和车辆方面显示出显着的性能改善,超过了人类监督的非模态方法。
Current methods for object detection, segmentation, and tracking fail in the presence of severe occlusions in busy urban environments. Labeled real data of occlusions is scarce (even in large datasets) and synthetic data leaves a domain gap, making it hard to explicitly model and learn occlusions. In this work, we present the best of both the real and synthetic worlds for automatic occlusion supervision using a large readily available source of data: time-lapse imagery from stationary webcams observing street intersections over weeks, months, or even years. We introduce a new dataset, Watch and Learn Time-lapse (WALT), consisting of 12 (4K and 1080p) cameras capturing urban environments over a year. We exploit this real data in a novel way to automatically mine a large set of unoccluded objects and then composite them in the same views to generate occlusions. This longitudinal self-supervision is strong enough for an amodal network to learn object-occluder-occluded layer representations. We show how to speed up the discovery of unoccluded objects and relate the confidence in this discovery to the rate and accuracy of training occluded objects. After watching and automatically learning for several days, this approach shows significant performance improvement in detecting and segmenting occluded people and vehicles, over human-supervised amodal approaches.