High Performance Embedded Architectures and Compilers - Fourth International Conference, HiPEAC 2009, Paphos, Cyprus, January 25-28, 2009. Proceedings

High Performance Embedded Architectures and Compilers - Fourth International Conference, HiPEAC 2009, Paphos, Cyprus, January 25-28, 2009. Proceedings
复制标题

高性能嵌入式架构和编译器 - 第四届国际会议,HiPEAC 2009,塞浦路斯帕福斯,2009 年 1 月 25-28 日。

DOI:
10.1007/978-3-540-92990-1_14
复制
发表时间:
2009
期刊:
--
影响因子:
--
通讯作者:
Howes L
Howes L
中科院分区:
--
文献类型:
--
作者:
Howes L

文献摘要

相似文献

在具有软件管理内存的多核架构上,有效地编排数据移动对性能至关重要,但繁琐且容易出错。在本文中,我们表明,当程序员可以明确指定的内存访问模式和计算内核的执行时间表,编译器或运行时系统可以得到有效的数据移动,即使内核代码的分析是困难的或不可能的。我们已经开发了一个框架的C++类解耦访问/执行规范,允许自动通信优化,如软件流水线和数据重用。我们证明了使用这些类编程的细胞宽带引擎架构的易用性和效率,通过实现一组基准,表现出数据重用和非仿射访问功能,并通过比较这些实现替代实现,其中使用手写DMA传输和基于软件的缓存。
On multi-core architectures with software-managed memories, effectively orchestrating data movement is essential to performance, but is tedious and error-prone. In this paper we show that when the programmer can explicitly specify both the memory access pattern and the execution schedule of a computation kernel, the compiler or run-time system can derive efficient data movement, even if analysis of kernel code is difficult or impossible. We have developed a framework of C++ classes for decoupled Access/Execute specifications, allowing for automatic communication optimisations such as software pipelining and data reuse. We demonstrate the ease and efficiency of programming the Cell Broadband Engine architecture using these classes by implementing a set of benchmarks, which exhibit data reuse and non-affine access functions, and by comparing these implementations against alternative implementations, which use hand-written DMA transfers and software-based caching.