Energy-efficient interconnect via Router Parking

Energy-efficient interconnect via Router Parking
复制标题

DOI:
10.1109/hpca.2013.6522345
复制
发表时间:
2013-02
期刊:
2013 IEEE 19th International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
A. Samih;Ren Wang;A. Krishna;C. Maciocco;T. Tai;Yan Solihin
A. Samih;Ren Wang;A. Krishna;C. Maciocco;T. Tai;Yan Solihin
中科院分区:
其他
文献类型:
--
作者:
A. Samih;Ren Wang;A. Krishna;C. Maciocco;T. Tai;Yan Solihin

文献摘要

被引文献

相似文献

芯片多处理器(CMPS)中片上核心计数的增加导致了网状和环状等互连的采用,这些互连消耗了越来越多的芯片功耗。此外,随着技术和电压的不断缩小,静态功率消耗在总功率中所占的比例越来越大;降低静态功率对能量比例计算越来越重要。目前,处理器设计者努力让未充分利用的内核进入深度休眠状态,以减少空闲功率并提高整体能效。然而,即使在最先进的CMP设计中,当核心进入休眠状态时,连接到它的路由器仍保持活动状态,以便继续转发数据包。在本文中,我们提出了路由器驻留-选择性功率门控路由器连接到驻留的核心。路由器驻留可确保保持网络连接,并限制数据包绕过驻留的路由器所造成的平均互连延迟影响。我们提出了两种路由器驻留算法--一种积极的方法来驻留尽可能多的路由器,另一种保守的方法是驻留有限的路由器,以便将对延迟增加的影响保持在最小。此外,我们提出了一种自适应策略,在运行时在两种算法之间进行选择。我们使用来自SPEC CPU2006和PARSEC 2.1基准测试套件的合成流量和实际工作负载来评估我们的算法。我们的评估结果显示,路由器驻留可以显著节省总互连能量(合成、SPEC CPU2006和Parsec 2.1工作负载的平均能耗分别为32%、40%和41%)。
The increase in on-chip core counts in Chip Multi Processors (CMPs) has led to the adoption of interconnects such as Mesh and Torus, which consume an increasing fraction of the chip power. Moreover, as technology and voltage continue to scale down, static power consumes a larger fraction of the total power; reducing it is increasingly important for energy proportional computing. Currently, processor designers strive to send under-utilized cores into deep sleep states in order to reduce idling power and improve overall energy efficiency. However, even in state-of-the-art CMP designs, when a core goes to sleep the router attached to it remains active in order to continue packet forwarding. In this paper, we propose Router Parking - selectively power-gating routers attached to parked cores. Router Parking ensures that network connectivity is maintained, and limits the average interconnect latency impact of packet detouring around parked routers. We present two Router Parking algorithms - an aggressive approach to park as many routers as possible, and a conservative approach that parks a limited set of routers in order to keep the impact on latency increase minimal. Further, we propose an adaptive policy to choose between the two algorithms at run-time. We evaluate our algorithms using both synthetic traffic as well as real workloads taken from SPEC CPU2006 and PARSEC 2.1 benchmark suites. Our evaluation results show that Router Parking can achieve significant savings in the total interconnect energy (average of 32%, 40% and 41% for the synthetic, SPEC CPU2006, and PARSEC 2.1 workloads, respectively).