Universality of Gradient Descent Neural Network Training

Universality of Gradient Descent Neural Network Training
复制标题

DOI:
10.1016/j.neunet.2022.02.016
复制
发表时间:
2020-07
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
G. Welper
G. Welper
中科院分区:
其他
文献类型:
--
作者:
G. Welper

文献摘要

相似文献

已经观察到,神经网络的设计选择对于它们的成功优化通常是至关重要的。因此,在这篇文章中,我们将讨论这样一个问题:重新设计神经网络是否总是可能的,以便它能够很好地利用梯度下降进行训练。这产生了如下普适性结果:对于给定的网络,如果有任何算法可以为分类任务找到良好的网络权重,则存在该网络的扩展,该扩展仅通过梯度下降训练来重现相同的前向模型。该结构不是用于实际计算,但它提供了一些关于元学习和相关方法中的预训练网络的可能性的方向。
It has been observed that design choices of neural networks are often crucial for their successful optimization. In this article, we therefore discuss the question if it is always possible to redesign a neural network so that it trains well with gradient descent. This yields the following universality result: If, for a given network, there is any algorithm that can find good network weights for a classification task, then there exists an extension of this network that reproduces the same forward model by mere gradient descent training. The construction is not intended for practical computations, but it provides some orientation on the possibilities of pre-trained networks in meta-learning and related approaches.