Automatic Locality Exploitation in the Codelet Model

Automatic Locality Exploitation in the Codelet Model
复制标题

DOI:
10.1109/trustcom.2013.104
复制
发表时间:
2013-07
期刊:
2013 12th IEEE International Conference on Trust, Security and Privacy in Computing and Communications
影响因子:
--
通讯作者:
Cheng Chen;Yao Wu;Joshua D. Suetterlein;Long Zheng;M. Guo;G. Gao
Cheng Chen;Yao Wu;Joshua D. Suetterlein;Long Zheng;M. Guo;G. Gao
中科院分区:
其他
文献类型:
--
作者:
Cheng Chen;Yao Wu;Joshua D. Suetterlein;Long Zheng;M. Guo;G. Gao

文献摘要

被引文献

相似文献

最先进的代码片段调度侧重于代码片段的动态工作负载平衡(类似于任务)。虽然由于计算资源得到充分利用,这种方法可以获得合理的性能,但可能无法达到最佳的节能效果。本文针对IBM Cyclops64多核系统,提出了一种新的多项式时间算法,该算法根据最大局域性和最小全局内存访问来找出最优的编码调度。我们的算法利用关于代码片段之间的位置的静态信息来实现更好的性能和能源效率。通过使用本地缓冲区将一个代码块中产生的数据传递给另一个代码块,可以大大减少全局内存访问。在我们开发的IBM Cyclops-64模拟器上的实验结果表明,与最先进的码块调度相比,我们的算法减少了高达59.7%的全局内存访问,实现了高达68.1%的性能改进,降低了高达40.7%的能耗。
State-of-the-art codelet scheduling focuses on dynamic workload balance of codelets (similar to tasks). While this approach may achieve reasonable performance since computation resources are fully utilized, it may not attain optimal energy savings. In this paper, targeting at IBM Cyclops64 -- a manycore system, we propose a novel polynomial time algorithm that finds out the optimal codelet scheduling in terms of maximum locality and minimum global memory accesses. Our algorithm leverages static information regarding locality among codelets to achieve better performance and energy efficiency. By using local buffers to pass data produced in one codelet to another, global memory accesses can be greatly reduced. The experimental results on our developed IBM Cyclops-64 emulator show that the codelet scheduling of our algorithm removes up to 59.7% of global memory accesses, achieves up to 68.1% of performance improvement, and reduces up to 40.7% of energy consumption comparing to the state-of-the-art codelet scheduling.