Deconstructing Distributions: A Pointwise Framework of Learning

Deconstructing Distributions: A Pointwise Framework of Learning
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Gal Kaplun;Nikhil Ghosh;S. Garg;B. Barak;Preetum Nakkiran
Gal Kaplun;Nikhil Ghosh;S. Garg;B. Barak;Preetum Nakkiran
中科院分区:
其他
文献类型:
--
作者:
Gal Kaplun;Nikhil Ghosh;S. Garg;B. Barak;Preetum Nakkiran

文献摘要

相似文献

在机器学习中,我们传统上评估单个模型的性能,并对一组测试输入进行平均。在这项工作中,我们提出了一种新的方法:我们衡量的性能时,评估一个$\textit{单输入点}$的模型集合。具体来说,我们研究了一个点的$\textit{profile}$:模型在测试分布上的平均性能与它们在这个单独点上的逐点性能之间的关系。我们发现,配置文件可以产生新的见解模型和数据的结构-在和分布。例如,我们凭经验表明,真实的数据分布由具有定性不同轮廓的点组成。一方面,在逐点和平均表现之间存在具有强相关性的“兼容“点。另一方面,也有弱相关甚至负相关的点:在这种情况下,提高整体模型准确性实际上会损害这些输入的性能。我们证明,这些实验观察与先前工作中提出的几个简化的学习模型的预测是不一致的。作为一个应用程序,我们使用配置文件来构建一个数据集,我们称之为CIFAR-10-NEG:CINIC-10的一个子集,这样对于标准模型,CIFAR-10-NEG的准确性与CIFAR-10测试的准确性呈负相关。这是第一次展示了一个完全颠倒“在线准确性”的OOD数据集(米勒,陶里,Raghunathan,佐川,Koh,Shankar,梁,卡蒙和施密特2021)
In machine learning, we traditionally evaluate the performance of a single model, averaged over a collection of test inputs. In this work, we propose a new approach: we measure the performance of a collection of models when evaluated on a $\textit{single input point}$. Specifically, we study a point's $\textit{profile}$: the relationship between models' average performance on the test distribution and their pointwise performance on this individual point. We find that profiles can yield new insights into the structure of both models and data -- in and out-of-distribution. For example, we empirically show that real data distributions consist of points with qualitatively different profiles. On one hand, there are"compatible"points with strong correlation between the pointwise and average performance. On the other hand, there are points with weak and even $\textit{negative}$ correlation: cases where improving overall model accuracy actually $\textit{hurts}$ performance on these inputs. We prove that these experimental observations are inconsistent with the predictions of several simplified models of learning proposed in prior work. As an application, we use profiles to construct a dataset we call CIFAR-10-NEG: a subset of CINIC-10 such that for standard models, accuracy on CIFAR-10-NEG is $\textit{negatively correlated}$ with accuracy on CIFAR-10 test. This illustrates, for the first time, an OOD dataset that completely inverts"accuracy-on-the-line"(Miller, Taori, Raghunathan, Sagawa, Koh, Shankar, Liang, Carmon, and Schmidt 2021)