Accelerating arithmetic kernels with coherent attached FPGA coprocessors

Accelerating arithmetic kernels with coherent attached FPGA coprocessors
复制标题

使用连贯的附加 FPGA 协处理器加速算术内核

DOI:
10.7873/date.2015.1123
复制
发表时间:
2015
期刊:
2015 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
C. Hagleitner
C. Hagleitner
中科院分区:
--
文献类型:
--
作者:
Heiner Giefers;R. Polig;C. Hagleitner

文献摘要

被引文献

相似文献

通过将已知未充分利用CPU的计算内核迁移到基于FPGA的协处理器,可以提高计算机系统的能量效率。与需要显式数据移动的传统基于I/O的协处理器相比,相干连接的加速器可以在与主机CPU相同的虚拟地址空间上操作。共享内存组织使广泛接受的编程模型,并有助于在通用计算系统中部署节能加速器。在本文中,我们研究了一个FFT加速器的FPGA附加通过相干加速器处理器接口(CAPI)的POWER8处理器。我们的研究结果表明,连贯的附加加速器优于基于设备驱动程序的方法在延迟方面。与在12核CPU上运行的优化并行软件FFT相比,硬件加速可将能效提高5倍,并将单线程性能提高2倍以上。我们的结论是,集成到异构编程框架,如OpenCL的CAPI将促进延迟关键操作,并将进一步提高混合系统的可编程性。
The energy efficiency of computer systems can be increased by migrating computational kernels that are known to under-utilize the CPU to an FPGA based coprocessor. In contrast to traditional I/O-based coprocessors that require explicit data movement, coherently attached accelerators can operate on the same virtual address space than the host CPU. A shared memory organization enables widely accepted programming models and helps to deploy energy efficient accelerators in general purpose computing systems. In this paper we study an FFT accelerator on FPGA attached via the Coherent Accelerator Processor Interface (CAPI) to a POWER8 processor. Our results show that the coherent attached accelerator outperforms device driver based approaches in terms of latency. Hardware acceleration delivers a 5× gain in energy efficiency compared to an optimized parallel software FFT running on a 12-core CPU and improves single thread performance by more than 2×. We conclude that the integration of CAPI into heterogeneous programming frameworks such as OpenCL will facilitate latency critical operations and will further enhance programmability of hybrid systems.