Exploring Many-Core Design Templates for FPGAs and ASICs

Exploring Many-Core Design Templates for FPGAs and ASICs
复制标题

探索 FPGA 和 ASIC 的众核设计模板

DOI:
10.1155/2012/439141
复制
发表时间:
2012
期刊:
Int. J. Reconfigurable Comput.
影响因子:
--
通讯作者:
J. Wawrzynek
J. Wawrzynek
中科院分区:
--
文献类型:
--
作者:
Ilia A. Lebedev;Christopher W. Fletcher;Shaoyi Cheng;James C. Martin;Austin Doupnik;D. Burke;Mingjie Lin;J. Wawrzynek

文献摘要

被引文献

相似文献

我们提出了一种高效的方法来硬件设计的基础上,用于实现计算绑定的应用程序表示在一个高层次的数据并行语言,如OpenCL的多核microarchitectural模板。该模板通过一系列高级参数(如互连拓扑或处理元件架构)在每个应用程序的基础上进行定制。这种方法的主要优点是,它(i)允许程序员通过高级编程语言定义的API来表达并行性,(ii)支持粗粒度多线程和细粒度线程,同时允许位级资源控制,(iii)减少了为不同算法或不同应用程序重新设计系统所需的工作。我们比较模板驱动的设计,全定制和可编程的方法,通过研究在几个候选平台的计算绑定的数据并行贝叶斯图推理算法的实现。具体来说,我们研究了一系列基于模板的FPGA和ASIC平台上的实现,并与全定制设计进行比较。在整个研究中,我们使用通用图形处理单元(GPGPU)的实现作为性能和面积基线。我们表明,我们的方法,类似于生产力的可编程方法,如GPGPU应用程序,产生的实现与性能接近的FPGA和ASIC平台上的全定制设计。
We present a highly productive approach to hardware design based on a many-coremicroarchitectural template used to implement compute-bound applications expressed in a high-level data-parallel language such as OpenCL. The template is customized on a per-application basis via a range of high-level parameters such as the interconnect topology or processing element architecture. The key benefits of this approach are that it (i) allows programmers to express parallelism through an API defined in a high-level programming language, (ii) supports coarse-grained multithreading and fine-grained threading while permitting bit-level resource control, and (iii) reduces the effort required to repurpose the systemfor different algorithms or different applications. We compare template-driven design to both full-custom and programmable approaches by studying implementations of a compute-bound data-parallel Bayesian graph inference algorithm across several candidate platforms. Specifically, we examine a range of template-based implementations on both FPGA and ASIC platforms and compare each against full custom designs. Throughout this study, we use a general-purpose graphics processing unit (GPGPU) implementation as a performance and area baseline. We show that our approach, similar in productivity to programmable approaches such as GPGPU applications, yields implementations with performance approaching that of full-custom designs on both FPGA and ASIC platforms.