Hardware-Enabled Efficient Data Processing With Tensor-Train Decomposition

Hardware-Enabled Efficient Data Processing With Tensor-Train Decomposition
复制标题

DOI:
10.1109/tcad.2021.3058317
复制
发表时间:
2021-02
影响因子:
2.9
通讯作者:
Zheng Qu;Bangyan Wang;Hengnu Chen;Jilan Lin;Ling Liang;Guoqi Li;Zheng Zhang;Yuan Xie
Zheng Qu;Bangyan Wang;Hengnu Chen;Jilan Lin;Ling Liang;Guoqi Li;Zheng Zhang;Yuan Xie
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zheng Qu;Bangyan Wang;Hengnu Chen;Jilan Lin;Ling Liang;Guoqi Li;Zheng Zhang;Yuan Xie

文献摘要

相似文献

近年来,张量计算已成为解决大数据分析、机器学习、医学图像和EDA问题的一种有前途的工具。为了减轻张量处理的内存和计算强度,人们广泛地采用分解技术,特别是张量序列分解(TTD)来压缩极高维张量数据。尽管TTD具有打破维度诅咒的潜力,但研究人员尚未充分利用其计算潜力,这主要是因为两个原因:1)由于每次TTD迭代中的奇异值分解(SVD)运算,执行TTD本身是耗时和耗能的;2)在某些应用中,如深度学习推理,通常需要额外的软硬件优化来处理获得的TT格式的数据。在本文中,我们将通过两种方法解决这些挑战。首先,我们提出了一种算法-硬件协同设计的定制架构,即TTD引擎来加速TTD。我们使用MRI图像压缩作为演示应用来说明所提出的加速器的有效性。其次,我们给出了一个案例研究,展示了TT格式数据处理的好处和使用TTD引擎的有效性。在案例研究中,我们使用TT方法来实现卷积运算,这对于TT格式的数据来说是困难和不平凡的。实验结果表明,与CPU实现相比,TTD引擎的平均加速比为14.9倍~36.9倍;与GPU基准相比,TTD引擎的加速比平均为4.1倍~9.9倍。与CPU和GPU相比,能效分别提高了至少14.4倍和5.4倍。此外,我们支持硬件的TT格式数据处理进一步提高了复杂操作和应用程序的执行效率。
In recent years, tensor computation has become a promising tool for solving big data analysis, machine learning, medical image, and EDA problems. To ease the memory and computation intensity of tensor processing, decomposition techniques, especially tensor-train decomposition (TTD), are widely adopted to compress the extremely high-dimensional tensor data. Despite TTD’s potential to break the curse of dimensionality, researchers have not yet leveraged its full computational potential, mainly because of two reasons: 1) executing TTD itself is time- and energy-consuming due to the singular value decomposition (SVD) operation inside each of TTD’s iteration and 2) additional software/hardware optimizations are often required to process the obtained TT-format data in certain applications such as deep learning inference. In this article, we address these challenges with two approaches. First, we propose an algorithm-hardware co-design with customized architecture, namely, TTD Engine to accelerate TTD. We use MRI image compression as a demo application to illustrate the efficacy of the proposed accelerator. Second, we present a case study demonstrating the benefit of TT-format data processing and the efficacy of using TTD Engine. In the case study, we use the TT approach to realize convolution operation, which is difficult and nontrivial for TT-format data. Experimental results show that, TTD Engine achieves, on average, $14.9 \times $ – $36.9 \times $ speedup over CPU implementations and $4.1\times $ – $9.9\times $ speedup compared to the GPU baseline. The energy efficiency is also improved by at least $14.4\times $ and $5.4\times $ over CPU and GPU, respectively. Moreover, our hardware-enabled TT-format data processing further leads to more efficient implementations of complicated operations and applications.