μC-States: Fine-grained GPU datapath power management

μC-States: Fine-grained GPU datapath power management
复制标题

DOI:
10.1145/2967938.2967941
复制
发表时间:
2016-09
期刊:
2016 International Conference on Parallel Architecture and Compilation Techniques (PACT)
影响因子:
--
通讯作者:
Onur Kayiran;Adwait Jog;Ashutosh Pattnaik;Rachata Ausavarungnirun;Xulong Tang;M. Kandemir;G. Loh;
Onur Kayiran;Adwait Jog;Ashutosh Pattnaik;Rachata Ausavarungnirun;Xulong Tang;M. Kandemir;G. Loh;
中科院分区:
其他
文献类型:
--
作者:
Onur Kayiran;Adwait Jog;Ashutosh Pattnaik;Rachata Ausavarungnirun;Xulong Tang;M. Kandemir;G. Loh;

文献摘要

被引文献

相似文献

为了提高图形处理单元(GPU)的性能,而不仅仅是增加核心数量,架构师最近采用了一种扩展方法:GPU核心的峰值吞吐量和单个功能正在迅速增加。GPU中的这种大核心趋势导致了各种挑战,包括更高的静态功耗以及大核心的数据路径组件的更低和不平衡的利用率。正如我们在本文中所示,随之而来的两个关键问题:(1)较低和不平衡的数据路径利用率可能会浪费功率,因为应用程序并不总是利用大核心数据路径的所有部分,以及(2)在某些情况下,由于每个大核心产生的更多内存请求导致内存系统争用,使用大核心可能会导致应用程序性能下降。本文介绍了一种新的分析数据路径组件利用率在大核心GPU的排队论原理的基础上。在此分析的基础上,我们为整个数据路径引入了一种细粒度的动态功耗和时钟门控机制,称为μC-States,其目的是通过关闭或下调数据路径组件来最大限度地降低功耗,这些组件不是运行应用程序性能的瓶颈。我们的实验评估表明,μC-States显著降低了大核心GPU的静态和动态功耗,同时还显著提高了受高内存系统争用影响的应用程序的性能。我们还表明,我们的数据路径组件利用率的分析可以指导调度和设计决策的GPU架构,包含异构的核心。
To improve the performance of Graphics Processing Units (GPUs) beyond simply increasing core count, architects are recently adopting a scale-up approach: the peak throughput and individual capabilities of the GPU cores are increasing rapidly. This big-core trend in GPUs leads to various challenges, including higher static power consumption and lower and imbalanced utilization of the datapath components of a big core. As we show in this paper, two key problems ensue: (1) the lower and imbalanced datapath utilization can waste power as an application does not always utilize all portions of the big core datapath, and (2) the use of big cores can lead to application performance degradation in some cases due to the higher memory system contention caused by the more memory requests generated by each big core. This paper introduces a new analysis of datapath component utilization in big-core GPUs based on queuing theory principles. Building on this analysis, we introduce a fine-grained dynamic power- and clock-gating mechanism for the entire datapath, called μC-States, which aims to minimize power consumption by turning off or tuning-down datapath components that are not bottlenecks for the performance of the running application. Our experimental evaluation demonstrates that μC-States significantly reduces both static and dynamic power consumption in a big-core GPU, while also significantly improving the performance of applications affected by high memory system contention. We also show that our analysis of datapath component utilization can guide scheduling and design decisions in a GPU architecture that contains heterogeneous cores.