MT-DMA: A DMA Controller Supporting Efficient Matrix Transposition for Digital Signal Processing

MT-DMA: A DMA Controller Supporting Efficient Matrix Transposition for Digital Signal Processing
复制标题

MT-DMA:支持数字信号处理高效矩阵转置的 DMA 控制器

DOI:
10.1109/access.2018.2889558
复制
发表时间:
2019
期刊:
影响因子:
3.9
通讯作者:
Zhiying Wang
Zhiying Wang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Sheng Ma;Yuanwu Lei;Libo Huang;Zhiying Wang

文献摘要

相似文献

矩阵转置在数字信号处理中起着至关重要的作用。然而,现有的矩阵转置实现具有显著的局限性。传统的设计使用加载和存储指令来完成矩阵转置。根据加载/存储单元的数量,这种设计通常在每个时钟周期转置多达一个矩阵元素。更严重的是,这种设计不能并行执行矩阵转置和数据计算。现代数字信号处理器将对矩阵转置的支持集成到直接存储器访问(DMA)控制器中;矩阵可以在数据移动期间转置。它允许并行执行矩阵转置和数据计算。然而,它的带宽利用率是有限的,它只能传输一个矩阵元素每个时钟周期。为了解决现有设计的局限性,我们提出了矩阵转置DMA(MT-DMA),以支持DMA控制器中的高效矩阵转置。它可以在每个时钟周期转置多个矩阵元素,以提高带宽利用率。与现有的设计相比,MT-DMA实现了最大23.9倍的性能提高微基准测试。它也更节能。由于MT-DMA有效地隐藏了数据计算背后的矩阵转置延迟,因此它的性能非常接近真实的应用的理想设计。
Matrix transposition plays a critical role in digital signal processing. However, the existing matrix transposition implementations have significant limitations. A traditional design uses load and store instructions to accomplish matrix transposition. Depending on the amount of load/store units, this design typically transposes up to one matrix element per clock cycle. More seriously, this design cannot perform matrix transposition and data calculations in parallel. Modern digital signal processors integrate the support for matrix transposition into the direct memory access (DMA) controller; the matrix can be transposed during data movements. It allows the parallel execution of matrix transposition and data calculations. Yet, its bandwidth utilization is limited; it can only transfer one matrix element per clock cycle. To address the limitations of the existing designs, we propose matrix transposition DMA (MT-DMA), to support efficient matrix transposition in DMA controllers. It can transpose multiple matrix elements per clock cycle to improve the bandwidth utilization. Compared with the existing designs, MT-DMA achieves a maximum 23.9 times performance improvement for micro-benchmarks. It is also more energy efficient. Since MT-DMA effectively hides the latency of matrix transposition behind data calculations, it performs very closely to an ideal design for real applications.