Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models

Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Beidi Chen;Tri Dao;Kaizhao Liang;Jiaming Yang;Zhao Song;A. Rudra;C. Ré
Beidi Chen;Tri Dao;Kaizhao Liang;Jiaming Yang;Zhao Song;A. Rudra;C. Ré
中科院分区:
其他
文献类型:
--
作者:
Beidi Chen;Tri Dao;Kaizhao Liang;Jiaming Yang;Zhao Song;A. Rudra;C. Ré

文献摘要

相似文献

过度参数化的神经网络泛化能力很好,但训练成本很高。理想情况下,人们希望在保持其泛化优势的同时降低其计算成本。稀疏模型训练是实现这一目标的一种简单而有前途的方法,但由于现有方法存在精度损失、训练运行时间缓慢或难以稀疏化所有模型组件等问题,因此仍然存在挑战。其核心问题是在一组离散的稀疏矩阵上搜索稀疏掩码是困难和昂贵的。为了解决这个问题,我们的主要见解是优化一个连续的超集稀疏矩阵与一个固定的结构,称为产品的蝴蝶矩阵。由于蝶形矩阵不是硬件有效的,我们提出了简单的蝶形(块和平面)的变种,以利用现代硬件。我们的方法(像素化蝴蝶)使用基于平坦块蝴蝶和低秩矩阵的简单固定稀疏模式来稀疏大多数网络层(例如,注意,MLP)。我们通过经验验证,Pixelated Butterfly比Butterfly快3倍,并加快了训练速度,以实现有利的准确性-效率权衡。在ImageNet分类和WikiText-103语言建模任务中,我们的稀疏模型的训练速度比密集的MLP-Mixer,Vision Transformer和GPT-2介质快2.5倍,并且准确性没有下降。
Overparameterized neural networks generalize well but are expensive to train. Ideally, one would like to reduce their computational cost while retaining their generalization benefits. Sparse model training is a simple and promising approach to achieve this, but there remain challenges as existing methods struggle with accuracy loss, slow training runtime, or difficulty in sparsifying all model components. The core problem is that searching for a sparsity mask over a discrete set of sparse matrices is difficult and expensive. To address this, our main insight is to optimize over a continuous superset of sparse matrices with a fixed structure known as products of butterfly matrices. As butterfly matrices are not hardware efficient, we propose simple variants of butterfly (block and flat) to take advantage of modern hardware. Our method (Pixelated Butterfly) uses a simple fixed sparsity pattern based on flat block butterfly and low-rank matrices to sparsify most network layers (e.g., attention, MLP). We empirically validate that Pixelated Butterfly is 3x faster than butterfly and speeds up training to achieve favorable accuracy--efficiency tradeoffs. On the ImageNet classification and WikiText-103 language modeling tasks, our sparse models train up to 2.5x faster than the dense MLP-Mixer, Vision Transformer, and GPT-2 medium with no drop in accuracy.