MBAG: A Scalable Mini-Block Adaptive Gradient Method for Deep Neural Networks

MBAG: A Scalable Mini-Block Adaptive Gradient Method for Deep Neural Networks
复制标题

DOI:
10.1109/bigdata55660.2022.10020262
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Jaewoo Lee
Jaewoo Lee
中科院分区:
其他
文献类型:
--
作者:
Jaewoo Lee

文献摘要

相似文献

预处理是一种被广泛应用于加速优化算法收敛的技术。最近提出的高效二阶算法(如KFAC)表明,利用损失函数的曲率信息对梯度进行预处理可以帮助实现更快的收敛速度。然而,由于计算和存储成本较高,它们在大规模深度学习中的实用性仍然受到限制。在这项工作中,我们提出了一种随机自适应梯度算法,称为最小块自适应梯度算法(MBAG),它解决了预条件矩阵计算中的计算挑战。为了减少每次迭代的成本,MBAG使用矩阵求逆引理解析地计算预条件矩阵的逆,然后使用迭代求解器近似求其平方根。此外,为了减轻存储需求,MBAG将模型参数划分为小尺寸的子集,并且仅计算与每个参数子集相关联的预条件器子块。这极大地提高了算法的可扩展性。使用真实数据集,将MBAG算法与常用的一阶和二阶算法在自动编码和分类任务上的性能进行了比较。
Preconditioning is a technique widely used to accelerate the convergence of optimization algorithms. Recently proposed efficient second-order algorithms (such as KFAC) showed that preconditioning the gradient using the curvature information of loss function can help achieve faster convergence. However, their practicality in large-scale deep learning is still limited due to the high computational and storage cost. In this work, we propose a stochastic adaptive gradient algorithm, called Mini-Block Adaptive Gradient (MBAG), that addresses those computational challenges in computing the preconditioning matrix. To reduce the per-iteration cost, MBAG analytically computes the inverse of preconditioning matrix using the matrix inversion lemma and then approximately finds its square root using an iterative solver. Further, to mitigate the storage requirement, MBAG partitions model parameters into subsets of small size and only computes sub-blocks of preconditioner associated with each subset of parameters. This greatly improves the scalability of the proposed algorithm. The performance of MBAG is compared to that of popular first- and second-order algorithms on auto-encoder and classification tasks using real datasets.