Gamma: leveraging Gustavson’s algorithm to accelerate sparse matrix multiplication

Gamma: leveraging Gustavson’s algorithm to accelerate sparse matrix multiplication
复制标题

DOI:
10.1145/3445814.3446702
复制
发表时间:
2021-04
期刊:
Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Guowei Zhang;Nithya Attaluri;J. Emer;Daniel Sánchez
Guowei Zhang;Nithya Attaluri;J. Emer;Daniel Sánchez
中科院分区:
其他
文献类型:
--
作者:
Guowei Zhang;Nithya Attaluri;J. Emer;Daniel Sánchez

文献摘要

被引文献

相似文献

稀疏矩阵-稀疏矩阵乘法 (spMspM) 是各种科学和机器学习应用的核心。 spMspM 在通用架构上效率较低,这使得加速器具有吸引力。然而,先前的 spMspM 加速器使用内积或外积数据流,这些数据流的输入或输出重用性较差,从而导致高流量和较差的性能。这些先前的加速器尚未探索古斯塔夫森算法,这是一种替代的 spMspM 数据流,它不会遇到这些问题,但具有先前加速器不支持的不规则内存访问模式。我们推出了 GAMMA,一种 spMspM 加速器,它使用 Gustavson 算法来解决先前工作的挑战。 GAMMA 使用专门的处理元件和简单的高基数合并来执行 spMspM 的计算,并并行执行许多合并以实现高吞吐量。 GAMMA 使用新颖的片上存储结构,结合了高速缓存和显式管理缓冲区的功能。该结构捕获了古斯塔夫森的不规则重用模式,并通过显式解耦的数据移动传输数千个并发稀疏纤维(即行或列的坐标和值列表)。 GAMMA 采用新的动态调度算法,尽管存在不规则性,但仍能实现高利用率。我们还提出了新的预处理算法,可提高 GAMMA 的效率和多功能性。因此,GAMMA 的性能比之前的加速器高出 gmean 2.1 倍,并将内存流量减少了 gmean 2.2 倍,最高可达 13 倍。
Sparse matrix-sparse matrix multiplication (spMspM) is at the heart of a wide range of scientific and machine learning applications. spMspM is inefficient on general-purpose architectures, making accelerators attractive. However, prior spMspM accelerators use inner- or outer-product dataflows that suffer poor input or output reuse, leading to high traffic and poor performance. These prior accelerators have not explored Gustavson's algorithm, an alternative spMspM dataflow that does not suffer from these problems but features irregular memory access patterns that prior accelerators do not support. We present GAMMA, an spMspM accelerator that uses Gustavson's algorithm to address the challenges of prior work. GAMMA performs spMspM's computation using specialized processing elements with simple high-radix mergers, and performs many merges in parallel to achieve high throughput. GAMMA uses a novel on-chip storage structure that combines features of both caches and explicitly managed buffers. This structure captures Gustavson's irregular reuse patterns and streams thousands of concurrent sparse fibers (i.e., lists of coordinates and values for rows or columns) with explicitly decoupled data movement. GAMMA features a new dynamic scheduling algorithm to achieve high utilization despite irregularity. We also present new preprocessing algorithms that boost GAMMA's efficiency and versatility. As a result, GAMMA outperforms prior accelerators by gmean 2.1x, and reduces memory traffic by gmean 2.2x and by up to 13x.