Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation

Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
复制标题

DOI:
10.1007/978-3-030-58545-7_40
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Liang-Chieh Chen;Raphael Gontijo Lopes;Bowen Cheng;Maxwell D. Collins;E. D. Cubuk;Barret Zoph;Hartwig Adam;J. Shlens
Liang-Chieh Chen;Raphael Gontijo Lopes;Bowen Cheng;Maxwell D. Collins;E. D. Cubuk;Barret Zoph;Hartwig Adam;J. Shlens
中科院分区:
其他
文献类型:
--
作者:
Liang-Chieh Chen;Raphael Gontijo Lopes;Bowen Cheng;Maxwell D. Collins;E. D. Cubuk;Barret Zoph;Hartwig Adam;J. Shlens

文献摘要

被引文献

相似文献

大型判别模型中的监督学习是现代计算机视觉的支柱。这种方法需要投资于大规模的人类注释数据集,以实现最先进的结果。反过来,监督学习的功效可能会受到人类注释数据集大小的限制。这种限制对于图像分割任务尤其明显,其中人类注释的费用特别大,但可能存在大量未标记的数据。在这项工作中,我们询问是否可以在未标记的视频序列和额外图像中利用半监督学习来提高城市场景分割的性能,同时解决语义,实例和全景分割。这项工作的目标是避免构建特定于标签传播的复杂的、学习过的架构(例如,片匹配和光流)。相反,我们只是为未标记的数据预测伪标签,并使用人工注释和伪标记的数据训练后续模型。该过程重复多次。因此,我们的Naive-Student模型,经过这种简单而有效的迭代半监督学习训练,在所有三个Cityscapes基准测试中都达到了最先进的结果,在测试集上达到了67.8% PQ,42.6% AP和85.2% mIOU的性能。我们认为这项工作是朝着建立一个简单的程序来利用未标记的视频序列和额外的图像,以超越核心计算机视觉任务的最先进性能迈出的重要一步。
Supervised learning in large discriminative models is a mainstay for modern computer vision. Such an approach necessitates investing in large-scale human-annotated datasets for achieving state-of-the-art results. In turn, the efficacy of supervised learning may be limited by the size of the human annotated dataset. This limitation is particularly notable for image segmentation tasks, where the expense of human annotation is especially large, yet large amounts of unlabeled data may exist. In this work, we ask if we may leverage semi-supervised learning in unlabeled video sequences and extra images to improve the performance on urban scene segmentation, simultaneously tackling semantic, instance, and panoptic segmentation. The goal of this work is to avoid the construction of sophisticated, learned architectures specific to label propagation (e.g., patch matching and optical flow). Instead, we simply predict pseudo-labels for the unlabeled data and train subsequent models with both human-annotated and pseudo-labeled data. The procedure is iterated for several times. As a result, our Naive-Student model, trained with such simple yet effective iterative semi-supervised learning, attains state-of-the-art results at all three Cityscapes benchmarks, reaching the performance of 67.8% PQ, 42.6% AP, and 85.2% mIOU on the test set. We view this work as a notable step towards building a simple procedure to harness unlabeled video sequences and extra images to surpass state-of-the-art performance on core computer vision tasks.