Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model

Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
复制标题

DOI:
10.48550/arxiv.2305.15265
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Zirui Liu;Guanchu Wang;Shaochen Zhong;Zhaozhuo Xu;D. Zha;Ruixiang Tang;Zhimeng Jiang;Kaixiong Zhou-Kaixio
Zirui Liu;Guanchu Wang;Shaochen Zhong;Zhaozhuo Xu;D. Zha;Ruixiang Tang;Zhimeng Jiang;Kaixiong Zhou-Kaixio
中科院分区:
其他
文献类型:
--
作者:
Zirui Liu;Guanchu Wang;Shaochen Zhong;Zhaozhuo Xu;D. Zha;Ruixiang Tang;Zhimeng Jiang;Kaixiong Zhou-Kaixio

文献摘要

相似文献

随着模型大小的快速增长,由于其大量的内存使用,对大型预训练语言模型进行微调变得越来越困难。以前的工作通常侧重于减少网络中可训练参数的数量。虽然模型参数确实会影响内存使用,但训练期间的主要内存瓶颈来自于存储特征图(也称为激活),因为它们对于梯度计算至关重要。值得注意的是,神经网络通常使用随机梯度下降进行训练。我们认为,在随机优化中,只要梯度估计器无偏且方差合理,模型就可以处理噪声梯度。遵循这个动机,我们提出了一个新的无偏估计家族,称为 WTA-CRS,用于减少方差的矩阵生成,只需要存储子采样激活来计算梯度。我们的工作提供了理论和实验证据,表明在调整变压器的背景下,我们提出的估计器与现有估计器相比表现出较低的方差。通过用 Transformer 中的近似线性运算替换线性运算,我们可以在几乎没有精度下降的情况下实现高达 2.7$\times$ 的峰值内存减少,并可实现高达 $6.4\times$ 的更大批量大小。在相同的硬件下,WTA-CRS 通过应用更大的模型和/或更大批量大小的更快训练速度来实现更好的下游任务性能。
With the rapid growth in model size, fine-tuning the large pre-trained language model has become increasingly difficult due to its extensive memory usage. Previous works usually focus on reducing the number of trainable parameters in the network. While the model parameters do contribute to memory usage, the primary memory bottleneck during training arises from storing feature maps, also known as activations, as they are crucial for gradient calculation. Notably, neural networks are usually trained using stochastic gradient descent. We argue that in stochastic optimization, models can handle noisy gradients as long as the gradient estimator is unbiased with reasonable variance. Following this motivation, we propose a new family of unbiased estimators called WTA-CRS, for matrix production with reduced variance, which only requires storing the sub-sampled activations for calculating the gradient. Our work provides both theoretical and experimental evidence that, in the context of tuning transformers, our proposed estimators exhibit lower variance compared to existing ones. By replacing the linear operation with our approximated one in transformers, we can achieve up to 2.7$\times$ peak memory reduction with almost no accuracy drop and enables up to $6.4\times$ larger batch size. Under the same hardware, WTA-CRS enables better down-streaming task performance by applying larger models and/or faster training speed with larger batch sizes.