Cascading structured pruning: enabling high data reuse for sparse DNN accelerators

Cascading structured pruning: enabling high data reuse for sparse DNN accelerators
复制标题

DOI:
10.1145/3470496.3527419
复制
发表时间:
2022-06
期刊:
Proceedings of the 49th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
Edward Hanson;Shiyu Li;H. Li;Yiran Chen
Edward Hanson;Shiyu Li;H. Li;Yiran Chen
中科院分区:
其他
文献类型:
--
作者:
Edward Hanson;Shiyu Li;H. Li;Yiran Chen

文献摘要

相似文献

运行现代深度神经网络 (DNN) 的性能和效率在很大程度上受到数据移动的限制。为了缓解数据移动瓶颈,最近的 DNN 推理加速器设计广泛采用积极的压缩技术和稀疏跳过机制。这些机制避免使用零值权重或激活进行传输或计算,以节省时间和精力。然而,这种稀疏跳过逻辑涉及大的输入缓冲区和不规则的数据访问模式,从而排除了许多节能数据重用机会和数据流。在这项工作中,我们提出了级联结构化修剪(CSP),这种技术可以保留更多的数据重用机会,从而提高能源效率,同时保持与 SparTen 等最新稀疏架构相当的性能。 CSP 包括以下两个组件: 在算法级别,CSP-A 引入可预测的稀疏模式,允许对权重数据进行低开销压缩以及对激活和权重数据的顺序访问。在架构级别,CSP-H 利用 CSP-A 的诱导稀疏模式和新颖的数据流,仅访问一次唯一的激活数据,从而消除了对大型输入缓冲区的需求。每个 CSP-H 处理元件 (PE) 均采用新颖的累积缓冲区设计和基于计数器的稀疏跳跃机制,以最小的控制器开销支持数据流。我们在几个代表性模型上验证了我们的方法。我们的模拟结果表明,在大多数评估下,CSP 的能效比 SparTen 平均提高了 15 倍,并且加速比相当或更高。
Performance and efficiency of running modern Deep Neural Networks (DNNs) are heavily bounded by data movement. To mitigate the data movement bottlenecks, recent DNN inference accelerator designs widely adopt aggressive compression techniques and sparse-skipping mechanisms. These mechanisms avoid transferring or computing with zero-valued weights or activations to save time and energy. However, such sparse-skipping logic involves large input buffers and irregular data access patterns, thus precluding many energy-efficient data reuse opportunities and dataflows. In this work, we propose Cascading Structured Pruning (CSP), a technique that preserves significantly more data reuse opportunities for higher energy efficiency while maintaining comparable performance relative to recent sparse architectures such as SparTen. CSP includes the following two components: At algorithm level, CSP-A induces a predictable sparsity pattern that allows for low-overhead compression of weight data and sequential access to both activation and weight data. At architecture level, CSP-H leverages CSP-A's induced sparsity pattern with a novel dataflow to access unique activation data only once, thus removing the demand for large input buffers. Each CSP-H processing element (PE) employs a novel accumulation buffer design and a counter-based sparse-skipping mechanism to support the dataflow with minimum controller overhead. We verify our approach on several representative models. Our simulated results show that CSP achieves on average 15× energy efficiency improvement over SparTen with comparable or superior speedup under most evaluations.