Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions

Laughing Hyena Distillery: Extracting Compact Recurrences From Convolutions
复制标题

DOI:
10.48550/arxiv.2310.18780
复制
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Stefano Massaroli;Michael Poli;Daniel Y. Fu;Hermann Kumbong;Rom N. Parnichkun;Aman Timalsina;David W. Romero;Quinn McIntyre;Beidi Chen;A. Rudra;Ce Zhang;Christopher Ré;Stefano Ermon;Y. Bengio
Stefano Massaroli;Michael Poli;Daniel Y. Fu;Hermann Kumbong;Rom N. Parnichkun;Aman Timalsina;David W. Romero;Quinn McIntyre;Beidi Chen;A. Rudra;Ce Zhang;Christopher Ré;Stefano Ermon;Y. Bengio
中科院分区:
其他
文献类型:
--
作者:
Stefano Massaroli;Michael Poli;Daniel Y. Fu;Hermann Kumbong;Rom N. Parnichkun;Aman Timalsina;David W. Romero;Quinn McIntyre;Beidi Chen;A. Rudra;Ce Zhang;Christopher Ré;Stefano Ermon;Y. Bengio

文献摘要

相似文献

无注意序列模型的最新进展依赖于卷积作为变形金刚核心的注意操作员的替代方案。特别是,长卷积序列模型已经在许多域中实现了最先进的性能,但是在自动回归推理工作负载中产生了巨大的成本 - 天真地要求在输入序列上进行完整的通过(或激活的缓存)每个产生的令牌 - 类似于基于注意的模型。在本文中,我们试图在任何预训练的长卷积体系结构中启用$ \ Mathcal O(1)$计算和记忆成本,以减少记忆足迹并增加发电期间的吞吐量。具体而言,我们的方法包括从每个卷积层中提取低维线性空间模型,并建立在有理插值和模型级还原技术的基础上。我们进一步引入了基于卷积的层(例如Hyena)的建筑改进:通过将跨通道的过滤器重量缩小到头部,我们实现了更高的训练质量,并减少要蒸馏的过滤器数量。所得模型的吞吐量比变形金刚高10倍,在1.3B参数下,吞吐量高1.5倍,蒸馏后质量上没有任何损失。
Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads -- naively requiring a full pass (or caching of activations) over the input sequence for each generated token -- similarly to attention-based models. In this paper, we seek to enable $\mathcal O(1)$ compute and memory cost per token in any pre-trained long convolution architecture to reduce memory footprint and increase throughput during generation. Concretely, our methods consist in extracting low-dimensional linear state-space models from each convolution layer, building upon rational interpolation and model-order reduction techniques. We further introduce architectural improvements to convolution-based layers such as Hyena: by weight-tying the filters across channels into heads, we achieve higher pre-training quality and reduce the number of filters to be distilled. The resulting model achieves 10x higher throughput than Transformers and 1.5x higher than Hyena at 1.3B parameters, without any loss in quality after distillation.