Accelerating Sparse MTTKRP for Tensor Decomposition on FPGA

Accelerating Sparse MTTKRP for Tensor Decomposition on FPGA
复制标题

加速 FPGA 上张量分解的稀疏 MTTKRP

DOI:
10.1145/3543622.3573179
复制
发表时间:
2023
期刊:
ACM/SIGDA International Symposium on Field Programmable Gate Arrays
影响因子:
--
通讯作者:
Prasanna, Viktor
Prasanna, Viktor
中科院分区:
--
文献类型:
--
作者:
Wijeratne, Sasindu;Wang, Ta-Yang;Kannan, Rajgopal;Prasanna, Viktor

文献摘要

参考文献

被引文献

相似文献

稀疏矩阵化张量时间 Khatri-Rao 产品 (spMTTKRP) 是稀疏张量分解中计算量最大的内核。在本文中,我们提出了一种 FPGA 上的硬件算法协同设计,以最小化 spMTTKRP 在输入张量的所有模式上的执行时间。我们引入了 FLYCOO,这是一种新颖的张量格式,可以在沿所有模式计算 spMTTKRP 期间消除中间值与 FPGA 外部存储器的通信。我们使用 FLYCOO 对张量进行重新映射还平衡了多个处理引擎 (PE) 之间的工作负载。我们提出了一种并行算法,可以同时处理彼此独立的输入张量的多个分区。所提出的算法还在运行时动态地对张量进行排序,以增加外部存储器访问的数据局部性。我们开发了一种定制 FPGA 加速器设计,其中 (1) PE 由一系列管道组成,可以同时处理输入张量的多个元素,(2) 内存控制器可利用计算的外部内存访问的空间和时间局部性。与广泛使用的现实世界稀疏张量数据集上最先进的 CPU 和 GPU 实现相比,我们的工作在执行时间上实现了 8.8 倍和 3.8 倍的几何平均加速。
Sparse Matricized Tensor Times Khatri-Rao Product (spMTTKRP) is the most computationally intensive kernel in sparse tensor decomposition. In this paper, we propose a hardware-algorithm co-design on FPGA to minimize the execution time of spMTTKRP along all modes of an input tensor. We introduce FLYCOO, a novel tensor format that eliminates the communication of intermediate values to the FPGA external memory during the computation of spMTTKRP along all the modes. Our remapping of the tensor using FLYCOO also balances the workload among multiple Processing Engines (PEs). We propose a parallel algorithm that can concurrently process multiple partitions of the input tensor independent of each other. The proposed algorithm also orders the tensor dynamically during runtime to increase the data locality of the external memory accesses. We develop a custom FPGA accelerator design with (1) PEs consisting of a collection of pipelines that can concurrently process multiple elements of the input tensor and (2) memory controllers to exploit the spatial and temporal locality of the external memory accesses of the computation. Our work achieves a geometric mean of 8.8X and 3.8X speedup in execution time compared with the state-of-the-art CPU and GPU implementations on widely-used real-world sparse tensor datasets.
用于稀疏矩阵化张量时间的可重构低延迟内存系统 FPGA 上的 Khatri-Rao 产品
DOI: 10.1109/hpec49654.2021.9622851
发表时间: 2021
期刊: 2021
影响因子: --
作者:
Wijeratne, Sasindu;Kannan, Rajgopal;Prasanna, Viktor
通讯作者: Prasanna, Viktor
DOI: 10.1145/3295500.3356216
发表时间: 2019-11
期刊: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
Israt Nisa;Jiajia Li;Aravind Sukumaran-Rajam;Prasant Singh;Sri-ram Krishnamoorthy;P. Sadayappan;Singh Rawat;Sri-ram Krishnamoorthy;An Efficient Mixed-Mode
通讯作者: Israt Nisa;Jiajia Li;Aravind Sukumaran-Rajam;Prasant Singh;Sri-ram Krishnamoorthy;P. Sadayappan;Singh Rawat;Sri-ram Krishnamoorthy;An Efficient Mixed-Mode
DOI: --
发表时间: 2020
期刊: IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子: --
作者:
Zhiyu Cheng;Baopu Li;Yanwen Fan;Sid Ying
通讯作者: Sid Ying