Online power-performance adaptation of multithreaded programs using hardware event-based prediction

Online power-performance adaptation of multithreaded programs using hardware event-based prediction
复制标题

DOI:
10.1145/1183401.1183426
复制
发表时间:
2006-06
期刊:
--
影响因子:
--
通讯作者:
Matthew Curtis-Maury;James Dzierwa;C. Antonopoulos;Dimitrios S. Nikolopoulos
Matthew Curtis-Maury;James Dzierwa;C. Antonopoulos;Dimitrios S. Nikolopoulos
中科院分区:
其他
文献类型:
--
作者:
Matthew Curtis-Maury;James Dzierwa;C. Antonopoulos;Dimitrios S. Nikolopoulos

文献摘要

被引文献

相似文献

对于具有多核/多线程处理器和高组件密度的高端系统,功耗感知型高性能多线程库成为系统软件堆栈的关键元素。从用户级运行时库中在线调整多线程代码的能力和性能是一个相对较新且尚未探索的研究领域。我们提出了一个用户级的库框架,用于低功耗,高性能执行的多线程代码的近最佳在线适应。我们的框架通过调节并发性和改变程序执行时的处理器/线程配置来运行。它的创新之处在于,它使用来自硬件事件驱动分析的快速运行时性能预测来选择实现接近最佳能效点的线程粒度。预测器的使用大大降低了粒度控制和程序自适应的运行时成本。我们的框架实现了性能和ED 2(能量延迟平方)水平:i)与Oracle导出的离线预测器相当或更好; ii)明显优于使用穷举或本地化线性搜索的在线预测器。完整的预测和自适应框架是在一个采用英特尔超线程处理器的真实的多SMT系统上实现的,并在OpenMP程序中嵌入了自适应功能。
With high-end systems featuring multicore/multithreaded processors and high component density, power-aware high-performance multithreading libraries become a critical element of the system software stack. Online power and performance adaptation of multithreaded code from within user-level runtime libraries is a relatively new and unexplored area of research. We present a user-level library framework for nearly optimal online adaptation of multithreaded codes for low-power, high-performance execution. Our framework operates by regulating concurrency and changing the processors/threads configuration as the program executes. It is innovative in that it uses fast, runtime performance prediction derived from hardware event-driven profiling, to select thread granularities that achieve nearly optimal energy-efficiency points. The use of predictors substantially reduces the runtime cost of granularity control and program adaptation. Our framework achieves performance and ED2 (energy-delay-squared) levels which are: i) comparable to or better than those of oracle-derived offline predictors; ii) significantly better than those of online predictors using exhaustive or localized linear search. The complete prediction and adaptation framework is implemented on a real multi-SMT system with Intel Hyperthreaded processors and embeds adaptation capabilities in OpenMP programs.