APOLLO: An Automated Power Modeling Framework for Runtime Power Introspection in High-Volume Commercial Microprocessors

APOLLO: An Automated Power Modeling Framework for Runtime Power Introspection in High-Volume Commercial Microprocessors
复制标题

DOI:
10.1145/3466752.3480064
复制
发表时间:
2021-10
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Zhiyao Xie;Xiaoqing Xu;Matt Walker;Joshua Knebel;K. Palaniswamy;Nicolas Hebert;Jiang Hu;Huanrui Yang;Yiran Chen;Shidhartha Das
Zhiyao Xie;Xiaoqing Xu;Matt Walker;Joshua Knebel;K. Palaniswamy;Nicolas Hebert;Jiang Hu;Huanrui Yang;Yiran Chen;Shidhartha Das
中科院分区:
其他
文献类型:
--
作者:
Zhiyao Xie;Xiaoqing Xu;Matt Walker;Joshua Knebel;K. Palaniswamy;Nicolas Hebert;Jiang Hu;Huanrui Yang;Yiran Chen;Shidhartha Das

文献摘要

被引文献

相似文献

准确的功耗建模对于节能 CPU 设计和运行时管理至关重要。理想的功率建模框架需要准确而快速,实现高时间分辨率(理想情况下是周期精确),同时运行时计算开销低,并且可以通过自动化轻松扩展到不同的设计。尽管之前进行了大量研究,但同时满足这些相互冲突的目标具有挑战性,并且基本上尚未实现。在本文中,我们提出了 APOLLO,这是一种自动化的每周期功耗建模框架,可作为设计时功耗估计器和低开销运行时片上功率计 (OPM) 的基础。 APOLLO 使用基于极小最大凹罚分 (MCP) 的特征选择算法,自动选择小于 0.05% 的 RTL 信号作为功率代理。功耗估算在 Arm Neoverse N1 [3] 上分别实现 R2 > 0.95,在 Arm Cortex-A77 [2] 微处理器上实现 R2 > 0.94。当与仿真器辅助流程集成时,APOLLO 可在几分钟内完成百万门工业 CPU 设计的数百万周期基准的每周期功耗估算。此外,功率模型被综合并集成到微处理器实现中作为运行时 OPM。当首选粗粒度时间分辨率时,APOLLO 的准确性进一步提高。据我们所知,这是第一个在不影响精度的情况下同时实现每周期时间分辨率和面积/功耗开销的运行时 OPM,并在高性能、无序工业 CPU 设计上得到了验证。
Accurate power modeling is crucial for energy-efficient CPU design and runtime management. An ideal power modeling framework needs to be accurate yet fast, achieve high temporal resolution (ideally cycle-accurate) yet with low runtime computational overheads, and easily extensible to diverse designs through automation. Simultaneously satisfying such conflicting objectives is challenging and largely unattained despite significant prior research. In this paper, we propose APOLLO, an automated per-cycle power modeling framework that serves as the basis for both a design-time power estimator and a low-overhead runtime on-chip power meter (OPM). APOLLO uses the minimax concave penalty (MCP)-based feature selection algorithm to automatically select less than 0.05% of RTL signals as power proxies. The power estimation achieves R2 > 0.95 on Arm Neoverse N1 [3] and R2 > 0.94 on Arm Cortex-A77 [2] microprocessors, respectively. When integrated with an emulator-assisted flow, APOLLO finishes per-cycle power estimation on millions-of-cycles benchmark in minutes for million-gate industrial CPU designs. Furthermore, the power model is synthesized and integrated into the microprocessor implementation as a runtime OPM. APOLLO’s accuracy further improves when coarse-grained temporal resolution is preferred. To our best knowledge, this is the first runtime OPM that simultaneously achieves per-cycle temporal resolution and area/power overhead without compromising accuracy, which is validated on high-performance, out-of-order industrial CPU designs.