Cascaded DMA Controller for Speedup of Indirect Memory Access in Irregular Applications

Cascaded DMA Controller for Speedup of Indirect Memory Access in Irregular Applications
复制标题

DOI:
10.1109/ia349570.2019.00017
复制
发表时间:
2019-11
期刊:
2019 IEEE/ACM 9th Workshop on Irregular Applications: Architectures and Algorithms (IA3)
影响因子:
--
通讯作者:
Tomoya Kashimata;T. Kitamura;K. Kimura;H. Kasahara
Tomoya Kashimata;T. Kitamura;K. Kimura;H. Kasahara
中科院分区:
其他
文献类型:
--
作者:
Tomoya Kashimata;T. Kitamura;K. Kimura;H. Kasahara

文献摘要

被引文献

相似文献

由稀疏线性代数计算引起的间接内存访问在重要的实际应用中有着广泛的应用。然而,它们也会导致严重的低效内存访问和流水线停顿,从而导致即使在高内存带宽和大量计算资源的情况下,执行效率也很低。间接存储器访问(例如访问A[B[i]])的一个重要问题是它需要两个后续的不同的存储器访问:索引加载(B[i])和以下数据元素访问(A[B[i]])。为了克服这种情况,我们提出了级联DMAC(CDMAC)。除了CPU核、矢量加速器和本地数据存储器外,该CDMAC还打算连接到多核芯片的每个核中。它在片外主存储器和向加速器提供数据的核内本地数据存储器之间执行数据传输。CDMAC的核心思想是将两个DMAC级联,使得第一个DMAC加载索引,第二个DMAC通过使用这些索引来访问数据元素。因此,该组织通过给出索引数组和元素数组来实现自主的间接内存访问,并通过将稀疏数据排入本地数据内存来获得高效的SIMD计算。我们在一块FPGA板上实现了一个具有所提出的CDMAC的多核处理器。稀疏矩阵向量乘法在FPGA上的测试结果表明,CDMAC与CPU数据传输相比,最多可获得17倍的加速比。
Indirect memory accesses caused by sparse linear algebra calculations are widely used in important real applications. However, they also cause serious inefficient memory accesses and pipeline stalls resulting low execution efficiency even with high memory bandwidth and much computational resource. One of the important issues of indirect memory accesses, such as accessing A[B[i]], is it requires two succeeding different memory accesses: the index loads (B[i]) and the following data element accesses (A[B[i]]). To overcome this situation, we propose the Cascaded-DMAC (CDMAC). This CDMAC is intended to be attached in each core of a multicore chip in addition to a CPU core, a vector accelerator, and a local data memory. It performs data transfers between an off-chip main memory and an in-core local data memory, which provides data to the accelerator. The key idea of the CDMAC is cascading two DMACs so that the first one loads indices, then the second one accesses data elements by using these indices. Thus, this organization realizes the autonomous indirect memory accesses by giving an index array and an element array, and obtains the efficient SIMD computations by lining up the sparse data into the local data memory. We implemented a multicore processor having the proposed CDMAC on an FPGA board. The evaluation result of sparse matrix-vector multiplications on the FPGA shows that the CDMAC achieves 17x speedup at most compared with the CPU data transfer.