Exploring FPGA-specific Optimizations for Irregular OpenCL Applications

Exploring FPGA-specific Optimizations for Irregular OpenCL Applications
复制标题

探索针对不规则 OpenCL 应用的 FPGA 特定优化

DOI:
10.1109/reconfig.2018.8641699
复制
发表时间:
2018
期刊:
2018 International Conference on ReConFigurable Computing and FPGAs (ReConFig)
影响因子:
--
通讯作者:
Y. Hanafy
Y. Hanafy
中科院分区:
--
文献类型:
--
作者:
Mohamed W. Hassan;A. Helal;P. Athanas;W. Feng;Y. Hanafy

文献摘要

被引文献

相似文献

OpenCL正在作为一种高级硬件说明语言出现,以应对开发FPGA应用程序的生产力挑战。与传统的硬件说明语言(HDLS)不同,OpenCL提供了一个抽象接口,以促进高生产率,使最终用户能够快速描述所需的计算,包括并行性和数据移动,以为其应用程序创建自定义的硬件加速器。但是,这些实现的加速器不太可能在不采用FPGA特异性优化的情况下有效地利用可重构结构,尤其是对于不规则的OpenCL应用程序。因此,我们探索了针对OPENCL应用程序的FPGA特异性优化空间,并提供了有关优化技术改善应用程序性能和资源利用率的见解。探索此优化空间将使最终用户能够利用FPGA的计算潜力。虽然这些优化是一般且适用于任何应用程序,但预期的绩效增益和资源利用效率取决于应用程序特征。具体而言,硬件剖道师用于分析OpenCL应用程序内核的局限性,并指导FPGA优化实现的开发。特别是,我们追求了不规则的OpenCL应用程序更具挑战性的问题,这些应用患有工作负载不平衡,不可预测的控制流以及不规则的记忆与访问模式。使用来自图形遍历,组合逻辑和稀疏线性代数应用域中的代表性内核的实验表明,FPGA特异性优化可以改善与OpenDEntection-agrogencent opencl Code相比,最多可以提高不规则OpenCL应用的性能基准套件。
OpenCL is emerging as a high-level hardware description language to address the productivity challenges of developing applications on FPGAs. Unlike traditional hardware description languages (HDLs), OpenCL provides an abstract interface to facilitate high productivity, enabling end users to rapidly describe the required computations, including parallelism and data movement, to create custom hardware accelerators for their applications. However, these OpenCL-realized accelerators are unlikely to make efficient use of the reconfigurable fabric without adopting FPGA-specific optimizations, particularly for irregular OpenCL applications. Consequently, we explore the FPGA-specific optimization space for OpenCL applications and present insights on which optimization techniques improve application performance and resource utilization. Exploring this optimization space will enable end users to harness the computational potential of the FPGA.While these optimizations are general and applicable to any application, the expected performance gain and resource-utilization efficiency vary depending on the application characteristics. Specifically, hardware profilers are used to analyze the limitations of OpenCL application kernels and to guide the development of FPGA-optimized implementations. In particular, we pursue the more challenging problem of irregular OpenCL applications, which suffer from workload imbalance, unpredictable control flow, and irregular memory-access patterns. Experiments using representative kernels from the graph traversal, combinational logic, and sparse linear algebra application domains show that FPGA-specific optimizations can improve the performance of irregular OpenCL applications by up to 27-fold in comparison to the architecture-agnostic OpenCL code from the OpenDwarfs benchmark suite.