When is memorization of irrelevant training data necessary for high-accuracy learning?

When is memorization of irrelevant training data necessary for high-accuracy learning?
复制标题

什么时候为了高精度学习需要记忆不相关的训练数据?

DOI:
10.1145/3406325.3451131
复制
发表时间:
2021
期刊:
ACM Symposium on the Theory of Computation (STOC
影响因子:
--
通讯作者:
Talwar, Kunal
Talwar, Kunal
中科院分区:
--
文献类型:
--
作者:
Brown, Gavin;Bun, Mark;Feldman, Vitaly;Smith, Adam;Talwar, Kunal

文献摘要

参考文献

被引文献

相似文献

现代机器学习模型非常复杂,并且经常对有关单个输入的惊人数量的信息进行编码。在极端情况下,复杂模型似乎会记住整个输入示例,包括看似无关的信息(例如,来自文本的社会安全号码)。在本文中,我们的目的是了解这种记忆是否是准确学习所必需的。我们描述自然预测问题,其中每个足够精确的训练算法必须在预测模型中编码基本上所有关于其训练示例的大子集的信息。即使当样本是高维的,并且熵比样本大小高得多,即使大多数信息最终与手头的任务无关,这仍然是正确的。此外,我们的结果不依赖于训练算法或用于学习的模型类别。我们的问题是下一个符号预测和聚类标记任务的简单且相当自然的变体。这些任务可以看作是文本和图像相关预测问题的抽象。为了建立我们的结果,我们减少从一个家庭的单向通信问题,我们证明了新的信息复杂性下界。
Modern machine learning models are complex and frequently encode surprising amounts of information about individual inputs. In extreme cases, complex models appear to memorize entire input examples, including seemingly irrelevant information (social security numbers from text, for example). In this paper, we aim to understand whether this sort of memorization is necessary for accurate learning. We describe natural prediction problems in which every sufficiently accurate training algorithm must encode, in the prediction model, essentially all the information about a large subset of its training examples. This remains true even when the examples are high-dimensional and have entropy much higher than the sample size, and even when most of that information is ultimately irrelevant to the task at hand. Further, our results do not depend on the training algorithm or the class of models used for learning.Our problems are simple and fairly natural variants of the next-symbol prediction and the cluster labeling tasks. These tasks can be seen as abstractions of text- and image-related prediction problems. To establish our results, we reduce from a family of one-way communication problems for which we prove new information complexity lower bounds.
通过通信复杂性对差异化私人学习进行样本复杂性限制
DOI: --
发表时间: 2014
期刊: SIAM journal on computing (Print)
影响因子: --
作者:
V. Feldman;David Xiao
通讯作者: David Xiao
PAC-贝叶斯框架的局限性
DOI: --
发表时间: 2020
期刊: Neural Information Processing Systems
影响因子: --
作者:
Roi Livni;S. Moran
通讯作者: S. Moran
DOI: --
发表时间: 2019-02
期刊: ArXiv
影响因子: --
作者:
Chiyuan Zhang;Samy Bengio;Moritz Hardt;Y. Singer
通讯作者: Chiyuan Zhang;Samy Bengio;Moritz Hardt;Y. Singer
描述纯私人学习者的样本复杂性
DOI: --
发表时间: 2019
影响因子: 6
作者:
A. Beimel;Kobbi Nissim;Uri Stemmer
通讯作者: Uri Stemmer
DOI: 10.2307/1426408
发表时间: 1969
影响因子: 1.2
作者:
J. Littlewood
通讯作者: J. Littlewood