Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient Descent

Lightweight Projective Derivative Codes for Compressed Asynchronous Gradient Descent
复制标题

DOI:
--
复制
发表时间:
2022-01
期刊:
--
影响因子:
--
通讯作者:
Pedro Soto;Ilia Ilmer;Haibin Guan;Jun Li
Pedro Soto;Ilia Ilmer;Haibin Guan;Jun Li
中科院分区:
其他
文献类型:
--
作者:
Pedro Soto;Ilia Ilmer;Haibin Guan;Jun Li

文献摘要

相似文献

编码分布式计算已成为在大型数据集上执行梯度下降以减轻离散和其他故障的常见做法。本文提出了一种新的算法,编码的偏导数本身,并进一步优化代码通过执行有损压缩的衍生码字通过最大化的码字中包含的信息,同时最小化的码字之间的信息。编码理论的这种应用的效用是在优化研究中观察到的事实的几何结果,即噪声在基于梯度下降的学习算法中是可容忍的,有时甚至是有帮助的,因为它有助于避免过拟合和局部最小值。这与当前许多关于分布式编码计算的传统工作形成鲜明对比,分布式编码计算的重点是从工作者恢复所有数据。第二个进一步的贡献是编码方案的低权重性质允许异步梯度更新,因为代码可以被迭代解码;即,工作者的任务可以立即被更新到更大的梯度中。方向导数总是方向向量的线性函数;因此,我们的框架是鲁棒的,因为它可以将线性编码技术应用于一般的机器学习框架,如深度神经网络。
Coded distributed computation has become common practice for performing gradient descent on large datasets to mitigate stragglers and other faults. This paper proposes a novel algorithm that encodes the partial derivatives themselves and furthermore optimizes the codes by performing lossy compression on the derivative codewords by maximizing the information contained in the codewords while minimizing the information between the codewords. The utility of this application of coding theory is a geometrical consequence of the observed fact in optimization research that noise is tolerable, sometimes even helpful, in gradient descent based learning algorithms since it helps avoid overfitting and local minima. This stands in contrast with much current conventional work on distributed coded computation which focuses on recovering all of the data from the workers. A second further contribution is that the low-weight nature of the coding scheme allows for asynchronous gradient updates since the code can be iteratively decoded; i.e., a worker's task can immediately be updated into the larger gradient. The directional derivative is always a linear function of the direction vectors; thus, our framework is robust since it can apply linear coding techniques to general machine learning frameworks such as deep neural networks.