Extracting Data Parallelism in Non-Stencil Kernel Computing by Optimally Coloring Folded Memory Conflict Graph

Extracting Data Parallelism in Non-Stencil Kernel Computing by Optimally Coloring Folded Memory Conflict Graph
复制标题

通过最佳着色折叠内存冲突图提取非模板内核计算中的数据并行性

DOI:
--
复制
发表时间:
2018
期刊:
Design Automation Conference
影响因子:
--
通讯作者:
Mingjie Lin
Mingjie Lin
中科院分区:
--
文献类型:
--
作者:
Juan Escobedo;Mingjie Lin

文献摘要

被引文献

相似文献

非模块内核计算中的不规则内存访问模式使众所周知的超平面 - [1],晶格[2]或基于底塞尔的[3] HLS技术无效。我们开发了一种优雅而有效的技术,该技术从高级软件代码中综合了内存最佳体系结构,以最大程度地提高应用程序特定的数据并行性。我们的基本思想是利用嵌入数据访问模式和计算结构中的图形结构,以执行最大化并行内存访问的内存库,同时既可以保存硬件和能耗。具体而言,我们优先于颜色一个加权冲突图,从折叠基本冲突图产生的加权冲突图以最大程度地减少内存冲突。最有趣的是,我们的基于图的方法可以使内存库数量之间的直接权衡并最大程度地减少内存冲突。我们在六个基因标准的设备上,通过六基准计算内核进行经验测试我们的方法,通过降低冲突来测量。特别是,我们的方法仅需要9.56%的LUT,3.2%的FF,2.5%的BRAM和总可用硬件资源的11.33%DSP,即可获得映射功能,该功能可在修改后的前进高斯淘汰内部使用4个降低90%的冲突。同时内存访问。
Irregular memory access pattern in non-stencil kernel computing renders the well-known hyperplane- [1], lattice- [2], or tessellation-based [3] HLS techniques ineffective. We develop an elegant yet effective technique that synthesizes memory-optimal architecture from high level software code in order to maximize application-specific data parallelism. Our basic idea is to exploit graph structures embedded in data access pattern and computation structure in order to perform the memory banking that maximizes parallel memory accesses while conserving both hardware and energy consumption. Specifically, we priority color a weighted conflict graph generated from folding the fundamental conflict graph to maximize memory conflict reduction. Most interestingly, our graph-based methodology enables a straightforward tradeoff between the number of memory banks and minimizing memory conflicts.We empirically test our methodology with Vivado HLx 2015.4 on a standard Kintex-7 device for six benchmark computing kernels by measuring conflict reduction. In particular, our approach only require 9.56% LUT, 3.2% FF, 2.5% BRAM, and 11.33% DSP of the total available hardware resource to obtain a mapping function that achieves a 90% conflict reduction on a modified forward Gaussian elimination Kernel with 4 simultaneous memory accesses.