VC dimension of partially quantized neural networks in the overparametrized regime

VC dimension of partially quantized neural networks in the overparametrized regime
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yutong Wang;C. Scott
Yutong Wang;C. Scott
中科院分区:
其他
文献类型:
--
作者:
Yutong Wang;C. Scott

文献摘要

相似文献

Vapnik-Chervonenkis (VC)理论至今无法解释过参数化神经网络的小泛化误差。事实上,VC理论在大型网络中的现有应用获得了VC维的上界,该上界与权重的数量成正比,并且对于很大一类网络,这些上界已知是紧的。在这项工作中,我们专注于一类部分量化的网络,我们称之为超平面排列神经网络(HANNs)。通过样本压缩分析,我们发现HANNs的VC维明显小于权重的数量,同时具有很高的表达能力。特别是,在过参数化状态下,对HANNs的经验风险最小化实现了使用Lipschitz后验类概率进行分类的极大极小率。我们进一步从经验上论证了汉斯的表达能力。在121个UCI数据集的面板上,过度参数化的HANNs与最先进的全精度模型的性能相匹配。
Vapnik-Chervonenkis (VC) theory has so far been unable to explain the small generalization error of overparametrized neural networks. Indeed, existing applications of VC theory to large networks obtain upper bounds on VC dimension that are proportional to the number of weights, and for a large class of networks, these upper bound are known to be tight. In this work, we focus on a class of partially quantized networks that we refer to as hyperplane arrangement neural networks (HANNs). Using a sample compression analysis, we show that HANNs can have VC dimension significantly smaller than the number of weights, while being highly expressive. In particular, empirical risk minimization over HANNs in the overparametrized regime achieves the minimax rate for classification with Lipschitz posterior class probability. We further demonstrate the expressivity of HANNs empirically. On a panel of 121 UCI datasets, overparametrized HANNs match the performance of state-of-the-art full-precision models.