Generalization bounds via distillation

Generalization bounds via distillation
复制标题

通过蒸馏的泛化界限

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Lan Wang
Lan Wang
中科院分区:
--
文献类型:
--
作者:
Daniel Hsu;Ziwei Ji;Matus Telgarsky;Lan Wang

文献摘要

被引文献

相似文献

本文从理论上研究了以下经验现象:给定一个具有较差泛化界的高复杂度网络,人们可以将其提炼成具有几乎相同预测但具有较低复杂度和较小泛化界的网络。主要贡献是一个分析,表明原始网络从它的蒸馏中继承了这个良好的泛化界限,假设使用了表现良好的数据增强。这个边界以抽象和具体的形式呈现,后者由约简技术补充,以处理具有卷积层,全连接层和跳过连接等特征的现代计算图。为了完善这个故事,还提出了一个(更宽松的)经典压缩一致收敛分析,以及在cifar和mnist上进行的各种实验,证明了原始网络与其蒸馏之间相似的泛化性能。
This paper theoretically investigates the following empirical phenomenon: given a high-complexity network with poor generalization bounds, one can distill it into a network with nearly identical predictions but low complexity and vastly smaller generalization bounds. The main contribution is an analysis showing that the original network inherits this good generalization bound from its distillation, assuming the use of well-behaved data augmentation. This bound is presented both in an abstract and in a concrete form, the latter complemented by a reduction technique to handle modern computation graphs featuring convolutional layers, fully-connected layers, and skip connections, to name a few. To round out the story, a (looser) classical uniform convergence analysis of compression is also presented, as well as a variety of experiments on cifar and mnist demonstrating similar generalization performance between the original network and its distillation.