Adaptive sparse tiling for sparse matrix multiplication

Adaptive sparse tiling for sparse matrix multiplication
复制标题

DOI:
10.1145/3293883.3295712
复制
发表时间:
2019-02
期刊:
Proceedings of the 24th Symposium on Principles and Practice of Parallel Programming
影响因子:
--
通讯作者:
Changwan Hong;Aravind Sukumaran-Rajam;Israt Nisa;Kunal Singh;P. Sadayappan
Changwan Hong;Aravind Sukumaran-Rajam;Israt Nisa;Kunal Singh;P. Sadayappan
中科院分区:
其他
文献类型:
--
作者:
Changwan Hong;Aravind Sukumaran-Rajam;Israt Nisa;Kunal Singh;P. Sadayappan

文献摘要

被引文献

相似文献

瓷砖是用于数据局部性优化的关键技术,可广泛用于用于多核/多核CPU和GPU的密集矩阵矩阵乘法的高性能实现。但是,稀疏矩阵乘法的不规则和基质依赖性数据访问模式使使用平铺以增强数据重复使用的挑战。在本文中,我们设计了一种自适应瓷砖策略,并将其应用以增强两个原语的性能:SPMM(稀疏矩阵和密集矩阵的产物)和SDDMM(采样致密的矩阵乘法)。与已诉诸于非标准稀疏矩阵表示以提高性能的研究相反,我们使用标准压缩稀疏行(CSR)表示,在该表示中,在其中进行了行内的重新排序以启用自适应平铺。使用稀疏套件收集的大量矩阵进行的实验评估表明,与当前可用的最新替代方案相比,性能改善了。
Tiling is a key technique for data locality optimization and is widely used in high-performance implementations of dense matrix-matrix multiplication for multicore/manycore CPUs and GPUs. However, the irregular and matrix-dependent data access pattern of sparse matrix multiplication makes it challenging to use tiling to enhance data reuse. In this paper, we devise an adaptive tiling strategy and apply it to enhance the performance of two primitives: SpMM (product of sparse matrix and dense matrix) and SDDMM (sampled dense-dense matrix multiplication). In contrast to studies that have resorted to non-standard sparse-matrix representations to enhance performance, we use the standard Compressed Sparse Row (CSR) representation, within which intra-row reordering is performed to enable adaptive tiling. Experimental evaluation using an extensive set of matrices from the Sparse Suite collection demonstrates significant performance improvement over currently available state-of-the-art alternatives.