Identifying the optimal energy-efficient operating points of parallel workloads

Identifying the optimal energy-efficient operating points of parallel workloads
复制标题

DOI:
10.1109/iccad.2011.6105393
复制
发表时间:
2011-11
期刊:
2011 IEEE/ACM International Conference on Computer-Aided Design (ICCAD)
影响因子:
--
通讯作者:
Ryan Cochran;Can Hankendi;A. Coskun;S. Reda
Ryan Cochran;Can Hankendi;A. Coskun;S. Reda
中科院分区:
其他
文献类型:
--
作者:
Ryan Cochran;Can Hankendi;A. Coskun;S. Reda

文献摘要

被引文献

相似文献

随着每个处理器的内核数量的增加,人们强烈希望开发并行工作负载以利用硬件并行性。与单线程应用程序相比,由于线程交互和资源停顿,并行工作负载的特征更加复杂。本文提出了一种准确且可扩展的方法,用于在一组优化多核处理器能效的目标函数和约束下确定并行工作负载运行时的最佳系统操作点(即线程数和 DVFS 设置)。使用为商业多核系统上的各种并行工作负载收集的广泛训练数据集,我们构建了多项逻辑回归(MLR)模型,该模型根据工作负载特征来估计最佳系统设置。我们使用 L1 正则化来自动确定能源优化的相关工作负载指标。在运行时,我们的技术以可忽略的开销确定最佳线程数和 DVFS 设置。我们的实验表明,我们的方法优于现有技术,决策准确性提高了 51%。这意味着能源性能运行平均提高了 10.6%,最高提高了 30.9%。随着潜在系统操作点数量的增加,我们的技术还表现出卓越的可扩展性。
As the number of cores per processor grows, there is a strong incentive to develop parallel workloads to take advantage of the hardware parallelism. In comparison to single-threaded applications, parallel workloads are more complex to characterize due to thread interactions and resource stalls. This paper presents an accurate and scalable method for determining the optimal system operating points (i.e., number of threads and DVFS settings) at runtime for parallel workloads under a set of objective functions and constraints that optimize for energy efficiency in multi-core processors. Using an extensive training data set gathered for a wide range of parallel workloads on a commercial multi-core system, we construct multinomial logistic regression (MLR) models that estimate the optimal system settings as a function of workload characteristics. We use L1-regularization to automatically determine the relevant workload metrics for energy optimization. At runtime, our technique determines the optimal number of threads and the DVFS setting with negligible overhead. Our experiments demonstrate that our method outperforms prior techniques with up to 51% improved decision accuracy. This translates to up to 10.6% average improvement in energy-performance operation, with a maximum improvement of 30.9%. Our technique also demonstrates superior scalability as the number of potential system operating points increases.