LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation

LoSparse: Structured Compression of Large Language Models based on Low-Rank and Sparse Approximation
复制标题

DOI:
10.48550/arxiv.2306.11222
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yixiao Li;Yifan Yu;Qingru Zhang;Chen Liang;Pengcheng He;Weizhu Chen;Tuo Zhao
Yixiao Li;Yifan Yu;Qingru Zhang;Chen Liang;Pengcheng He;Weizhu Chen;Tuo Zhao
中科院分区:
其他
文献类型:
--
作者:
Yixiao Li;Yifan Yu;Qingru Zhang;Chen Liang;Pengcheng He;Weizhu Chen;Tuo Zhao

文献摘要

被引文献

相似文献

变形金刚在各种自然语言任务中取得了显着的结果,但是它们通常很大,需要大量的记忆和计算资源。为了降低这些模型的大小和复杂性,我们提出了losparse(低级和稀疏近似),这是一种新型的模型压缩技术,该技术通过低级别矩阵和稀疏矩阵的总和近似重量矩阵。我们的方法结合了低级别近似值和修剪的优势,同时避免了它们的局限性。低级别近似压缩神经元中的相干和表达部位,而修剪则消除了神经元中的不连贯和非表达部分。修剪会增强低级别近似值的多样性,而低级别近似可防止修剪失去太多表达神经元。我们评估了我们的自然语言理解,问题答案和自然语言生成任务的方法。我们表明,它的表现明显优于现有的压缩方法。
Transformer models have achieved remarkable results in various natural language tasks, but they are often prohibitively large, requiring massive memories and computational resources. To reduce the size and complexity of these models, we propose LoSparse (Low-Rank and Sparse approximation), a novel model compression technique that approximates a weight matrix by the sum of a low-rank matrix and a sparse matrix. Our method combines the advantages of both low-rank approximations and pruning, while avoiding their limitations. Low-rank approximation compresses the coherent and expressive parts in neurons, while pruning removes the incoherent and non-expressive parts in neurons. Pruning enhances the diversity of low-rank approximations, and low-rank approximation prevents pruning from losing too many expressive neurons. We evaluate our method on natural language understanding, question answering, and natural language generation tasks. We show that it significantly outperforms existing compression methods.