Exploiting mixed-mode parallelism for matrix operations on the HERA architecture through reconfiguration

Exploiting mixed-mode parallelism for matrix operations on the HERA architecture through reconfiguration
复制标题

通过重新配置在 HERA 架构上利用混合模式并行性进行矩阵运算

DOI:
10.1049/ip-cdt:20045136
复制
发表时间:
2006
期刊:
22nd International Conference on Field Programmable Logic and Applications (FPL)
影响因子:
--
通讯作者:
Sotirios G. Ziavras
Sotirios G. Ziavras
中科院分区:
--
文献类型:
--
作者:
Xiaofang Wang;Sotirios G. Ziavras

文献摘要

被引文献

相似文献

数百万门平台现场可编程门阵列 (FPGA) 的最新进展使得在可编程芯片上设计和实现复杂的并行系统成为可能,该芯片还包含硬件浮点单元 (FPU)。这些选项利用资源重新配置。与大多数仍然采用可重构逻辑来开发算法专用电路的 FPGA 社区相比,我们基于 FPGA 的混合模式可重构计算机可以同时实现多种并行执行模式,并且也是用户可编程的。我们的异构可重构架构(HERA)机器可以实现单指令、多数据(SIMD)、多指令、多数据(MIMD)和多SIMD(M-SIMD)执行模式。每个处理元件 (PE) 以具有紧密耦合本地内存的单精度 IEEE 754 FPU 为中心,并支持运行时在 SIMD 和 MIMD 之间动态切换。混合模式并行性有可能最好地匹配应用程序中所有子任务的特征,从而实现持续的高性能。 HERA 的性能通过两个常见的计算密集型测试平台进行评估:矩阵-矩阵乘法 (MMM) 和稀疏双边框块对角 (DBBD) 矩阵的 LU 分解。电力网络矩阵的实验结果表明,与 SIMD 和 MIMD 实现相比,LU 分解的混合模式调度可分别带来约 19% 和 15.5% 的加速。
Recent advances in multi-million-gate platform field-programmable gate arrays (FPGAs) have made it possible to design and implement complex parallel systems on a programmable chip that also incorporate hardware floating-point units (FPUs). These options take advantage of resource reconfiguration. In contrast to the majority of the FPGA community that still employs reconfigurable logic to develop algorithm-specific circuitry, our FPGA-based mixed-mode reconfigurable computing machine can implement simultaneously a variety of parallel execution modes and is also user programmable. Our heterogeneous reconfigurable architecture (HERA) machine can implement the single-instruction, multiple-data (SIMD), multiple-instruction, multiple-data (MIMD) and multiple-SIMD (M-SIMD) execution modes. Each processing element (PE) is centred on a single-precision IEEE 754 FPU with tightly-coupled local memory, and supports dynamic switching between SIMD and MIMD at runtime. Mixed-mode parallelism has the potential to best match the characteristics of all subtasks in applications, thus resulting in sustained high performance. HERA's performance is evaluated by two common computation-intensive testbenches: matrix-matrix multiplication (MMM) and LU factorisation of sparse doubly-bordered-block-diagonal (DBBD) matrices. Experimental results with electrical power network matrices show that the mixed-mode scheduling for LU factorisation can result in speedups of about 19% and 15.5% compared to the SIMD and MIMD implementations, respectively.