Universal halting times in optimization and machine learning

Universal halting times in optimization and machine learning
复制标题

优化和机器学习中的通用停止时间

DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Yann LeCun
Yann LeCun
中科院分区:
--
文献类型:
--
作者:
Levent Sagun;T. Trogdon;Yann LeCun

文献摘要

被引文献

相似文献

给出了两个随机系统:自旋眼镜系统和深度学习系统的优化算法停止时间的经验分布。给定一个算法,我们将其视为随机景观的优化例程和形式,停止时间的波动遵循一个分布,该分布在居中和缩放之后保持不变,即使当景观上的分布改变时也是如此。我们观察到两类定性的分布:在Google搜索、人类决策时间、QR特征值算法和自旋眼镜中出现的类Gumbel分布,以及在共轭梯度法、具有MNIST输入数据的深度网络和具有随机输入数据的深度网络中出现的类高斯分布。这一经验证据表明,在某些条件下,存在一类停顿时间独立于基础分布的分布。
The authors present empirical distributions for the halting time (measured by the number of iterations to reach a given accuracy) of optimization algorithms applied to two random systems: spin glasses and deep learning. Given an algorithm, which we take to be both the optimization routine and the form of the random landscape, the fluctuations of the halting time follow a distribution that, after centering and scaling, remains unchanged even when the distribution on the landscape is changed. We observe two qualitative classes: A Gumbel-like distribution that appears in Google searches, human decision times, the QR eigenvalue algorithm and spin glasses, and a Gaussian-like distribution that appears in conjugate gradient method, deep network with MNIST input data and deep network with random input data. This empirical evidence suggests presence of a class of distributions for which the halting time is independent of the underlying distribution under some conditions.