Cuttlefish: Low-Rank Model Training without All the Tuning

Cuttlefish: Low-Rank Model Training without All the Tuning
复制标题

DOI:
10.48550/arxiv.2305.02538
复制
发表时间:
2023-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Hongyi Wang;Saurabh Agarwal;Pongsakorn U-chupala;Yoshiki Tanaka;Eric P. Xing;Dimitris Papailiopoulos
Hongyi Wang;Saurabh Agarwal;Pongsakorn U-chupala;Yoshiki Tanaka;Eric P. Xing;Dimitris Papailiopoulos
中科院分区:
其他
文献类型:
--
作者:
Hongyi Wang;Saurabh Agarwal;Pongsakorn U-chupala;Yoshiki Tanaka;Eric P. Xing;Dimitris Papailiopoulos

文献摘要

相似文献

最近的研究表明,训练低秩神经网络可以有效地减少可训练参数的总数,而不会牺牲预测精度,从而实现端到端的加速。然而,低秩模型训练需要调整几个额外的因子分解超参数,例如每层因子分解的秩。在本文中,我们通过引入Cuttlefish来解决这一挑战,Cuttlefish是一种自动化的低秩训练方法,无需调整因子分解超参数。墨鱼利用了这样的观察,即经过几个时期的满秩训练,稳定的秩(即,真实秩的近似值)稳定在恒定值。一旦所有层的稳定秩收敛,Cuttlefish就从满秩训练切换到低秩训练,将每个分解的维度设置为其相应的稳定秩。我们的研究结果表明,Cuttlefish生成的模型比满秩模型小5.6倍,端到端训练过程快1.2倍,同时保持相当的准确性。此外,Cuttlefish的性能优于最先进的低秩模型训练方法和其他突出的基线。我们实现的源代码可以在https://github.com/hwang595/Cuttlefish上找到。
Recent research has shown that training low-rank neural networks can effectively reduce the total number of trainable parameters without sacrificing predictive accuracy, resulting in end-to-end speedups. However, low-rank model training necessitates adjusting several additional factorization hyperparameters, such as the rank of the factorization at each layer. In this paper, we tackle this challenge by introducing Cuttlefish, an automated low-rank training approach that eliminates the need for tuning factorization hyperparameters. Cuttlefish leverages the observation that after a few epochs of full-rank training, the stable rank (i.e., an approximation of the true rank) of each layer stabilizes at a constant value. Cuttlefish switches from full-rank to low-rank training once the stable ranks of all layers have converged, setting the dimension of each factorization to its corresponding stable rank. Our results show that Cuttlefish generates models up to 5.6 times smaller than full-rank models, and attains up to a 1.2 times faster end-to-end training process while preserving comparable accuracy. Moreover, Cuttlefish outperforms state-of-the-art low-rank model training methods and other prominent baselines. The source code for our implementation can be found at: https://github.com/hwang595/Cuttlefish.