MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts

MetaShift: A Dataset of Datasets for Evaluating Contextual Distribution Shifts and Training Conflicts
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Weixin Liang;James Y. Zou
Weixin Liang;James Y. Zou
中科院分区:
其他
文献类型:
--
作者:
Weixin Liang;James Y. Zou

文献摘要

相似文献

了解机器学习模型在不同数据分布中的性能对于可靠的应用至关重要。受此启发,人们越来越关注策划捕获分布变化的基准数据集。虽然有价值,但现有的基准是有限的,因为其中许多基准只包含少量的转变,而且它们缺乏关于不同转变之间差异的系统注释。我们提出MetaShift--一个包含410个类别的12,868组自然图像的集合--来解决这一挑战。我们利用Visual Genome及其注释的自然异质性来构建MetaShift。关键的构建思想是使用其元数据对图像进行聚类,元数据为每个图像提供上下文(例如,"有汽车的猫"或"浴室里的猫")表示不同的数据分布。MetaShift有两个重要的好处:第一,它包含比以前更多的自然数据转移数量级。其次,它提供了关于每个数据集的独特性的明确解释,以及衡量任何两个数据集之间分布变化量的距离分数。我们展示了MetaShift在基准测试中的实用性,这些基准测试最近提出了一些训练模型对数据变化具有鲁棒性的建议。我们发现,简单的经验风险最小化执行最好的变化时,是温和的,没有方法有系统的优势,大的变化。我们还展示了MetaShift如何帮助在模型训练期间可视化数据子集之间的冲突。
Understanding the performance of machine learning models across diverse data distributions is critically important for reliable applications. Motivated by this, there is a growing focus on curating benchmark datasets that capture distribution shifts. While valuable, the existing benchmarks are limited in that many of them only contain a small number of shifts and they lack systematic annotation about what is different across different shifts. We present MetaShift--a collection of 12,868 sets of natural images across 410 classes--to address this challenge. We leverage the natural heterogeneity of Visual Genome and its annotations to construct MetaShift. The key construction idea is to cluster images using its metadata, which provides context for each image (e.g."cats with cars"or"cats in bathroom") that represent distinct data distributions. MetaShift has two important benefits: first, it contains orders of magnitude more natural data shifts than previously available. Second, it provides explicit explanations of what is unique about each of its data sets and a distance score that measures the amount of distribution shift between any two of its data sets. We demonstrate the utility of MetaShift in benchmarking several recent proposals for training models to be robust to data shifts. We find that the simple empirical risk minimization performs the best when shifts are moderate and no method had a systematic advantage for large shifts. We also show how MetaShift can help to visualize conflicts between data subsets during model training.