Is Self-Supervised Learning More Robust Than Supervised Learning?

Is Self-Supervised Learning More Robust Than Supervised Learning?
复制标题

DOI:
10.48550/arxiv.2206.05259
复制
发表时间:
2022-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuanyi Zhong;Haoran Tang;Jun-Kun Chen;Jian Peng;Yu-Xiong Wang
Yuanyi Zhong;Haoran Tang;Jun-Kun Chen;Jian Peng;Yu-Xiong Wang
中科院分区:
其他
文献类型:
--
作者:
Yuanyi Zhong;Haoran Tang;Jun-Kun Chen;Jian Peng;Yu-Xiong Wang

文献摘要

相似文献

自我监督对比学习是学习无标签视觉表征的有力工具。以前的工作主要集中在评估各种预训练算法的识别精度上,而忽略了其他行为方面。除了准确性之外,分布稳健性对机器学习模型的可靠性也起着至关重要的作用。我们设计并进行了一系列稳健性测试,以量化对比学习和监督学习对下游或训练前数据分布变化的行为差异。这些测试利用多个级别的数据损坏,从像素级伽马失真到补丁级洗牌,再到数据集级别的分布偏移。我们的测试揭示了对比学习和监督学习有趣的稳健性行为。一方面,在下游腐败情况下,我们通常观察到对比学习比监督学习更稳健。另一方面,在训练前的污染情况下,我们发现对比学习容易受到补丁洗牌和像素强度变化的影响,但对数据集级别的分布变化不那么敏感。我们试图通过数据增强和特征空间属性的作用来解释这些结果。我们的见解对提高监督学习的下游稳健性具有一定的启示意义。
Self-supervised contrastive learning is a powerful tool to learn visual representation without labels. Prior work has primarily focused on evaluating the recognition accuracy of various pre-training algorithms, but has overlooked other behavioral aspects. In addition to accuracy, distributional robustness plays a critical role in the reliability of machine learning models. We design and conduct a series of robustness tests to quantify the behavioral differences between contrastive learning and supervised learning to downstream or pre-training data distribution changes. These tests leverage data corruptions at multiple levels, ranging from pixel-level gamma distortion to patch-level shuffling and to dataset-level distribution shift. Our tests unveil intriguing robustness behaviors of contrastive and supervised learning. On the one hand, under downstream corruptions, we generally observe that contrastive learning is surprisingly more robust than supervised learning. On the other hand, under pre-training corruptions, we find contrastive learning vulnerable to patch shuffling and pixel intensity change, yet less sensitive to dataset-level distribution change. We attempt to explain these results through the role of data augmentation and feature space properties. Our insight has implications in improving the downstream robustness of supervised learning.