Channel-Wise Early Stopping without a Validation Set via NNK Polytope Interpolation

Channel-Wise Early Stopping without a Validation Set via NNK Polytope Interpolation
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
2021 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
David Bonet;Antonio Ortega;Javier Ruiz-Hidalgo;Sarath Shekkizhar
David Bonet;Antonio Ortega;Javier Ruiz-Hidalgo;Sarath Shekkizhar
中科院分区:
其他
文献类型:
--
作者:
David Bonet;Antonio Ortega;Javier Ruiz-Hidalgo;Sarath Shekkizhar

文献摘要

被引文献

相似文献

最先进的神经网络架构继续在规模上扩展,并提供令人印象深刻的泛化结果,尽管这是以有限的可解释性为代价的。特别是,一个关键的挑战是确定何时停止训练模型,因为这对泛化有重大影响。卷积神经网络(ConvNets)包括由多个通道聚合形成的高维特征空间,由于维数灾难,分析中间数据表示和模型的演变可能具有挑战性。我们提出了通道式DeepNNK(CW-DeepNNK),这是一种基于非负核回归(NNK)图的新型通道式泛化估计,我们使用它在低维通道上执行局部多面体插值。这种方法导致基于实例的可解释性的学习数据表示和通道之间的关系。受我们观察的启发,我们使用CW-DeepNNK提出了一种新的早期停止标准,该标准(i)不需要验证集,(ii)基于任务性能指标,(iii)允许在每个通道的不同点停止。我们的实验表明,我们提出的方法相比,基于验证集性能的标准标准具有优势。
State-of-the-art neural network architectures continue to scale in size and deliver impressive generalization results, although this comes at the expense of limited interpretability. In particular, a key challenge is to determine when to stop training the model, as this has a significant impact on generalization. Convolutional neural networks (ConvNets) comprise high-dimensional feature spaces formed by the aggregation of multiple channels, where analyzing intermediate data representations and the model's evolution can be challenging owing to the curse of dimensionality. We present channel-wise DeepNNK (CW-DeepNNK), a novel channel-wise generalization estimate based on non-negative kernel regression (NNK) graphs with which we perform local polytope interpolation on low-dimensional channels. This method leads to instance-based interpretability of both the learned data representations and the relationship between channels. Motivated by our observations, we use CW-DeepNNK to propose a novel early stopping criterion that (i) does not require a validation set, (ii) is based on a task performance metric, and (iii) allows stopping to be reached at different points for each channel. Our experiments demonstrate that our proposed method has advantages as compared to the standard criterion based on validation set performance.