Identity Crisis: Memorization and Generalization under Extreme Overparameterization

Identity Crisis: Memorization and Generalization under Extreme Overparameterization
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Chiyuan Zhang;Samy Bengio;Moritz Hardt;Y. Singer
Chiyuan Zhang;Samy Bengio;Moritz Hardt;Y. Singer
中科院分区:
其他
文献类型:
--
作者:
Chiyuan Zhang;Samy Bengio;Moritz Hardt;Y. Singer

文献摘要

被引文献

相似文献

我们研究了在单个训练样本和身份映射任务的极端情况下,过参数网络的记忆和泛化之间的相互作用。我们研究了完全连通和卷积网络(FCN和CNN),无论是线性的还是非线性的,随机初始化,然后训练以最小化重建误差。训练过的网络通常采取两种形式之一:常量函数(记忆)和身份函数(泛化)。我们形式化地刻画了单层FCN和CNN中的泛化。我们的经验表明,不同的体系结构表现出显著不同的归纳偏向。例如,多达10层的CNN能够从单个示例中进行泛化,而FCN不能从60k个示例中可靠地学习身份函数。更深层的CNN经常失败,但仍然在记忆训练输出方面做了惊人的工作:因为CNN的偏差是位置不变的,所以模型必须通过许多层的协调从图像边界逐步增长输出模式。我们的工作有助于量化和可视化归纳偏差对体系结构选择的敏感性,例如深度、内核宽度和通道数量。
We study the interplay between memorization and generalization of overparameterized networks in the extreme case of a single training example and an identity-mapping task. We examine fully-connected and convolutional networks (FCN and CNN), both linear and nonlinear, initialized randomly and then trained to minimize the reconstruction error. The trained networks stereotypically take one of two forms: the constant function (memorization) and the identity function (generalization). We formally characterize generalization in single-layer FCNs and CNNs. We show empirically that different architectures exhibit strikingly different inductive biases. For example, CNNs of up to 10 layers are able to generalize from a single example, whereas FCNs cannot learn the identity function reliably from 60k examples. Deeper CNNs often fail, but nonetheless do astonishing work to memorize the training output: because CNN biases are location invariant, the model must progressively grow an output pattern from the image boundaries via the coordination of many layers. Our work helps to quantify and visualize the sensitivity of inductive biases to architectural choices such as depth, kernel width, and number of channels.