Enhancing Design Space Exploration by Extending CPU/GPU Specifications onto FPGAs

Enhancing Design Space Exploration by Extending CPU/GPU Specifications onto FPGAs
复制标题

DOI:
10.1145/2656207
复制
发表时间:
2015-03-01
影响因子:
2
通讯作者:
Ienne, Paolo
Ienne, Paolo
中科院分区:
计算机科学3区
文献类型:
--
作者:
Owaida, Muhsen;Falcao, Gabriel;Ienne, Paolo

文献摘要

被引文献

相似文献

复杂的专用计算系统的设计周期是极其昂贵和耗时的。它涉及多参数设计空间探索优化,其次是设计验证。特殊用途的VLSI实现的设计人员经常需要通过耗时的蒙特卡罗模拟来探索参数,例如最佳位宽和数据表示。这种基于模拟的探索过程的一个突出例子是用于纠错系统的解码器的设计,例如现代通信标准所采用的低密度奇偶校验(LDPC)码,其涉及每个设计点的数千次蒙特卡罗运行。目前,高性能计算提供了广泛的加速选项,从多核CPU到图形处理单元(GPU)和现场可编程门阵列(FPGA)。利用不同的目标体系结构通常与开发多个代码版本相关联,通常使用不同的编程范式。在这种情况下,我们评估了将单个OpenCL程序重新定位到多个平台的概念,从而显着减少设计时间。在多核CPU、GPU和FPGA上使用单个基于OpenCL的并行内核,无需修改或代码调优。我们使用SOpenCL(Silicon to OpenCL),这是一种自动将OpenCL内核转换为RTL的工具,以引入FPGA作为有效执行OpenCL中编码的模拟的潜在平台。我们使用LDPC解码仿真作为案例研究。通过测试范围从短/中(例如,8,000比特)到长长度(例如,64,800位)DVB-S2代码。我们观察到,取决于要模拟的设计参数,取决于设计的维度和阶段,GPU或FPGA可以更方便地适合不同的目的,从而提供与传统多核CPU不同的加速因子。
The design cycle for complex special-purpose computing systems is extremely costly and time-consuming. It involves a multiparametric design space exploration for optimization, followed by design verification. Designers of special purpose VLSI implementations often need to explore parameters, such as optimal bitwidth and data representation, through time-consuming Monte Carlo simulations. A prominent example of this simulation-based exploration process is the design of decoders for error correcting systems, such as the Low-Density Parity-Check (LDPC) codes adopted by modern communication standards, which involves thousands of Monte Carlo runs for each design point. Currently, high-performance computing offers a wide set of acceleration options that range from multicore CPUs to Graphics Processing Units (GPUs) and Field Programmable Gate Arrays (FPGAs). The exploitation of diverse target architectures is typically associated with developing multiple code versions, often using distinct programming paradigms. In this context, we evaluate the concept of retargeting a single OpenCL program to multiple platforms, thereby significantly reducing design time. A single OpenCL-based parallel kernel is used without modifications or code tuning on multicore CPUs, GPUs, and FPGAs. We use SOpenCL (Silicon to OpenCL), a tool that automatically converts OpenCL kernels to RTL in order to introduce FPGAs as a potential platform to efficiently execute simulations coded in OpenCL. We use LDPC decoding simulations as a case study. Experimental results were obtained by testing a variety of regular and irregular LDPC codes that range from short/medium (e.g., 8,000 bit) to long length (e.g., 64,800 bit) DVB-S2 codes. We observe that, depending on the design parameters to be simulated, on the dimension and phase of the design, the GPU or FPGA may suit different purposes more conveniently, thus providing different acceleration factors over conventional multicore CPUs.