Dynamic Voltage and Frequency Scaling in NoCs with Supervised and Reinforcement Learning Techniques

Dynamic Voltage and Frequency Scaling in NoCs with Supervised and Reinforcement Learning Techniques
复制标题

DOI:
10.1109/tc.2018.2875476
复制
发表时间:
2019-03-01
影响因子:
3.7
通讯作者:
Louri, Ahmed
Louri, Ahmed
中科院分区:
计算机科学2区
文献类型:
--
作者:
Fettes, Quintin;Clark, Mark;Louri, Ahmed

文献摘要

被引文献

相似文献

片上网络 (NoC) 因其规律性、效率、简单性和可扩展性而成为多核芯片中设计互连结构的实际选择。然而,由于晶体管漏电流以及内核与缓存之间的数据移动,NoC 遭受过多的静态功耗和动态能量的困扰。技术尺寸的不断缩小只会加剧功耗问题。动态电压和频率调节(DVFS)是一种旨在减少动态能量的技术;然而,这通常会以牺牲性能为代价。在本文中,我们提出了使用监督学习和强化学习方法的多核架构的 LEAD Learning 支持的能量感知动态电压/频率缩放。 LEAD 将路由器及其传出链路分组到同一 V/F 域中,并实施主动 DVFS 模式管理策略,该策略依赖于离线训练的机器学习模型,以便在不同电压/频率对之间提供最佳的 V/F 模式选择。我们提出了 LEAD 的三个监督学习版本,它们基于缓冲区利用率、缓冲区利用率的变化和能量/吞吐量的变化,允许基于对未来网络参数的准确预测进行主动模式选择。然后,我们描述了一种针对 LEAD 的强化学习方法,该方法直接优化 DVFS 模式选择,从而无需标签和阈值工程。在 4 x 4 集中式网格架构上使用 PARSEC 和 Splash-2 基准进行的仿真结果表明,通过使用监督学习 LEAD 可以实现平均动态节能 15.4%,吞吐量损失为 0.8%,并且对延迟没有显着影响。使用强化学习时,LEAD 将平均动态节能提高到 20.3%,但代价是吞吐量下降 1.5%,延迟增加 1.7%。总体而言,更灵活的强化学习方法能够在任何所需的能量与吞吐量权衡下学习更广泛的负载环境的最佳行为。
Network-on-Chips (NoCs) are the de facto choice for designing the interconnect fabric in multicore chips due to their regularity, efficiency, simplicity, and scalability. However, NoC suffers from excessive static power and dynamic energy due to transistor leakage current and data movement between the cores and caches. Power consumption issues are only exacerbated by ever decreasing technology sizes. Dynamic Voltage and Frequency Scaling (DVFS) is one technique that seeks to reduce dynamic energy; however this often occurs at the expense of performance. In this paper, we propose LEAD Learning-enabled Energy-Aware Dynamic voltage/frequency scaling for multicore architectures using both supervised learning and reinforcement learning approaches. LEAD groups the router and its outgoing links into the same V/F domain and implements proactive DVFS mode management strategies that rely on offline trained machine learning models in order to provide optimal V/F mode selection between different voltage/frequency pairs. We present three supervised learning versions of LEAD that are based on buffer utilization, change in buffer utilization and change in energy/throughput, which allow proactive mode selection based on accurate prediction of future network parameters. We then describe a reinforcement learning approach to LEAD that optimizes the DVFS mode selection directly, obviating the need for label and threshold engineering. Simulation results using PARSEC and Splash-2 benchmarks on a 4 x 4 concentrated mesh architecture show that by using supervised learning LEAD can achieve an average dynamic energy savings of 15.4 percent for a loss in throughput of 0.8 percent with no significant impact on latency. When reinforcement learning is used, LEAD increases average dynamic energy savings to 20.3 percent at the cost of a 1.5 percent decrease in throughput and a 1.7 percent increase in latency. Overall, the more flexible reinforcement learning approach enables learning an optimal behavior for a wider range of load environments under any desired energy versus throughput tradeoff.