Extending the WILDS Benchmark for Unsupervised Adaptation

Extending the WILDS Benchmark for Unsupervised Adaptation
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Shiori Sagawa;Pang Wei Koh;Tony Lee;Irena Gao;Sang Michael Xie;Kendrick Shen;Ananya Kumar;Weihua Hu
Shiori Sagawa;Pang Wei Koh;Tony Lee;Irena Gao;Sang Michael Xie;Kendrick Shen;Ananya Kumar;Weihua Hu
中科院分区:
其他
文献类型:
--
作者:
Shiori Sagawa;Pang Wei Koh;Tony Lee;Irena Gao;Sang Michael Xie;Kendrick Shen;Ananya Kumar;Weihua Hu

文献摘要

被引文献

相似文献

野外部署的机器学习系统通常在源分布上进行训练,但部署在不同的目标分布上。未标记数据可以成为缓解这些分布变化的强大杠杆,因为它通常比标记数据更可用,并且通常也可以从源分布之外的分布中获得。然而,现有的未标记数据的分布转移基准并不能反映现实应用中出现的场景的广度。在这项工作中,我们提出了 Wilds 2.0 更新,它扩展了 Wilds 分布变化基准中 10 个数据集中的 8 个,以包含在部署中实际可以获得的精选未标记数据。这些数据集涵盖广泛的应用(从组织学到野生动物保护)、任务(分类、回归和检测)和模式(照片、卫星图像、显微镜幻灯片、文本、分子图)。该更新通过使用相同的标记训练、验证和测试集以及评估指标来保持与原始 Wilds 基准的一致性。在这些数据集上,我们系统地对利用未标记数据的最先进方法进行了基准测试,包括域不变、自我训练和自我监督方法,并表明它们在 Wilds 上的成功是有限的。为了促进方法开发和评估,我们提供了一个开源包,可自动加载数据并包含本文中使用的所有模型架构和方法。代码和排行榜可在 https://wilds.stanford.edu 获取。
Machine learning systems deployed in the wild are often trained on a source distribution but deployed on a different target distribution. Unlabeled data can be a powerful point of leverage for mitigating these distribution shifts, as it is frequently much more available than labeled data and can often be obtained from distributions beyond the source distribution as well. However, existing distribution shift benchmarks with unlabeled data do not reflect the breadth of scenarios that arise in real-world applications. In this work, we present the Wilds 2.0 update, which extends 8 of the 10 datasets in the Wilds benchmark of distribution shifts to include curated unlabeled data that would be realistically obtainable in deployment. These datasets span a wide range of applications (from histology to wildlife conservation), tasks (classification, regression, and detection), and modalities (photos, satellite images, microscope slides, text, molecular graphs). The update maintains consistency with the original Wilds benchmark by using identical labeled training, validation, and test sets, as well as the evaluation metrics. On these datasets, we systematically benchmark state-of-the-art methods that leverage unlabeled data, including domain-invariant, self-training, and self-supervised methods, and show that their success on Wilds is limited. To facilitate method development and evaluation, we provide an open-source package that automates data loading and contains all of the model architectures and methods used in this paper. Code and leaderboards are available at https://wilds.stanford.edu .