SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing

SmartFRZ: An Efficient Training Framework using Attention-Based Layer Freezing
复制标题

DOI:
10.48550/arxiv.2401.16720
复制
发表时间:
2024-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Sheng Li;Geng Yuan;Yuezhen Dai;Youtao Zhang;Yanzhi Wang;Xulong Tang
Sheng Li;Geng Yuan;Yuezhen Dai;Youtao Zhang;Yanzhi Wang;Xulong Tang
中科院分区:
其他
文献类型:
--
作者:
Sheng Li;Geng Yuan;Yuezhen Dai;Youtao Zhang;Yanzhi Wang;Xulong Tang

文献摘要

相似文献

人工智能应用程序不断涌现,其中模型训练是为这些应用程序提供高质量服务的关键。然而,模型训练过程既耗时又耗能,不可避免地影响了用户对应用效率的需求。层冻结是一种有效的模型训练技术,为了提高训练效率而被提出。尽管现有的层冻结方法显示出降低模型训练成本的巨大潜力,但它们仍然存在缺乏通用性和准确性受损等缺点。例如,现有的层冻结方法要么需要在训练前手动定义冻结配置,这不适用于不同的网络,要么使用启发式冻结标准,很难保证在不同场景下的良好准确性。因此,缺乏一种通用且智能的层冻结方法,可以在训练过程中自动对不同网络进行“原位”层冻结。为此,我们提出了一个通用且高效的培训框架(SmartFRZ)。 SmartFRZ提出的核心技术是注意力引导层冻结,它可以在不影响准确性的情况下自动选择合适的层进行冻结。实验结果表明,SmartFRZ有效减少了训练计算量,实现了显着的训练加速,并且优于最先进的层冻结方法。
There has been a proliferation of artificial intelligence applications, where model training is key to promising high-quality services for these applications. However, the model training process is both time-intensive and energy-intensive, inevitably affecting the user's demand for application efficiency. Layer freezing, an efficient model training technique, has been proposed to improve training efficiency. Although existing layer freezing methods demonstrate the great potential to reduce model training costs, they still remain shortcomings such as lacking generalizability and compromised accuracy. For instance, existing layer freezing methods either require the freeze configurations to be manually defined before training, which does not apply to different networks, or use heuristic freezing criteria that is hard to guarantee decent accuracy in different scenarios. Therefore, there lacks a generic and smart layer freezing method that can automatically perform ``in-situation'' layer freezing for different networks during training processes. To this end, we propose a generic and efficient training framework (SmartFRZ). The core proposed technique in SmartFRZ is attention-guided layer freezing, which can automatically select the appropriate layers to freeze without compromising accuracy. Experimental results show that SmartFRZ effectively reduces the amount of computation in training and achieves significant training acceleration, and outperforms the state-of-the-art layer freezing approaches.