A Rule for Gradient Estimator Selection, with an Application to Variational Inference

A Rule for Gradient Estimator Selection, with an Application to Variational Inference
复制标题

梯度估计器选择规则及其在变分推理中的应用

DOI:
--
复制
发表时间:
2019
期刊:
International Conference on Artificial Intelligence and Statistics
影响因子:
--
通讯作者:
Justin Domke
Justin Domke
中科院分区:
--
文献类型:
--
作者:
Tomas Geffner;Justin Domke

文献摘要

被引文献

相似文献

随机梯度下降(SGD)是现代机器学习的主力。有时,可以使用许多不同的势梯度估计器。在这种情况下,选择成本和差异之间最好的折衷方案是很重要的。本文分析了SGD的收敛率作为时间的函数,而不是迭代的函数。这产生了一个简单的规则来选择导致最佳优化收敛保证的估计量。这种选择对于SGD的不同变体是相同的,并且具有关于目标的不同假设(例如,凹凸性或平滑性)。受这一原理的启发,我们提出了一种在给定有限估计量池时自动选择估计量的技术。然后,我们扩展到无限估计池,其中每个估计池都由控制变量权重索引。这是通过简化为混合整数二次规划实现的。根据经验,自动选择一个估计器的性能与事后选择的最佳估计器相当。
Stochastic gradient descent (SGD) is the workhorse of modern machine learning. Sometimes, there are many different potential gradient estimators that can be used. When so, choosing the one with the best tradeoff between cost and variance is important. This paper analyzes the convergence rates of SGD as a function of time, rather than iterations. This results in a simple rule to select the estimator that leads to the best optimization convergence guarantee. This choice is the same for different variants of SGD, and with different assumptions about the objective (e.g. convexity or smoothness). Inspired by this principle, we propose a technique to automatically select an estimator when a finite pool of estimators is given. Then, we extend to infinite pools of estimators, where each one is indexed by control variate weights. This is enabled by a reduction to a mixed-integer quadratic program. Empirically, automatically choosing an estimator performs comparably to the best estimator chosen with hindsight.