NTK-approximating MLP Fusion for Efficient Language Model Fine-tuning

NTK-approximating MLP Fusion for Efficient Language Model Fine-tuning
复制标题

DOI:
10.48550/arxiv.2307.08941
复制
发表时间:
2023-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianxin Wei;Zeming Guo;Yifan Chen;Jingrui He
Tianxin Wei;Zeming Guo;Yifan Chen;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Tianxin Wei;Zeming Guo;Yifan Chen;Jingrui He

文献摘要

相似文献

在许多自然语言处理应用中,微调预训练语言模型(PLM)已成为主要策略。然而,即使是对PLM进行微调以及进行推理也是昂贵的,特别是在计算能力较低的边缘设备上。一些通用方法(例如量化和蒸馏)已被广泛研究以减少PLM微调的计算量/内存,但很少有一次性压缩技术被探索。在本文中,我们研究了神经正切核(NTK)——它揭示了神经网络的梯度下降动态——在PLM中的多层感知器(MLP)模块,并提议通过NTK近似的MLP融合来创建一个轻量级的PLM。为了实现这一目标,我们将MLP重新视为一组子MLP,并将它们聚类为给定数量的质心,然后这些质心可以恢复为一个压缩的MLP,并且令人惊讶地表明它能很好地近似原始PLM的NTK。我们提供了在自然语言理解(NLU)和生成(NLG)任务上对PLM微调的大量实验,以验证所提出的MLP融合方法的有效性。我们的代码可在https://github.com/weitianxin/MLP_Fusion获取。
Fine-tuning a pre-trained language model (PLM) emerges as the predominant strategy in many natural language processing applications. However, even fine-tuning the PLMs and doing inference are expensive, especially on edge devices with low computing power. Some general approaches (e.g. quantization and distillation) have been widely studied to reduce the compute/memory of PLM fine-tuning, while very few one-shot compression techniques are explored. In this paper, we investigate the neural tangent kernel (NTK)--which reveals the gradient descent dynamics of neural networks--of the multilayer perceptrons (MLP) modules in a PLM and propose to coin a lightweight PLM through NTK-approximating MLP fusion. To achieve this, we reconsider the MLP as a bundle of sub-MLPs, and cluster them into a given number of centroids, which can then be restored as a compressed MLP and surprisingly shown to well approximate the NTK of the original PLM. Extensive experiments of PLM fine-tuning on both natural language understanding (NLU) and generation (NLG) tasks are provided to verify the effectiveness of the proposed method MLP fusion. Our code is available at https://github.com/weitianxin/MLP_Fusion.