Post-Silicon CPU Adaptation Made Practical Using Machine Learning

Post-Silicon CPU Adaptation Made Practical Using Machine Learning
复制标题

使用机器学习使硅后 CPU 适应变得实用

DOI:
10.1145/3307650.3322267
复制
发表时间:
2019
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Hong Wang
Hong Wang
中科院分区:
--
文献类型:
--
作者:
Stephen J. Tarsa;Rangeen Basu Roy Chowdhury;Julien Sébot;G. Chinya;Jayesh Gaur;K. Sankaranarayanan;Chit;R. Chappell;Ronak Singhal;Hong Wang

文献摘要

被引文献

相似文献

在运行时使架构适应工作负载的处理器有望实现引人注目的性能功耗比(PPW)增益,从而提供了一种减轻管道扩展带来的收益递减的方法。最先进的自适应CPU在片上部署机器学习(ML)模型,通过识别事件计数器数据中的工作负载模式来优化硬件。然而,尽管突破性的PPW增益,这样的设计尚未被广泛采用,由于在该领域的系统适应误差的潜力。本文介绍了一种基于Intel SkyLake的自适应CPU,它(1)关闭了部署的循环,(2)提供了一种新的硅后定制机制。我们的CPU执行预测集群门控,动态设置集群架构的问题宽度,同时时钟门控未使用的资源。门控决策由在现有微控制器上执行的ML自适应模型驱动,最大限度地降低了设计复杂性,并允许通过固件更新轻松调整性能特性。至关重要的是,我们表明,尽管适应模型可能会受到统计盲点的影响,可能会降低新工作负载的性能,但通过仔细的设计和训练,这些影响可以减少到最小。在SPEC2017上,我们的自适应CPU将PPW提高了31.4%,并且比最先进的非自适应CPU少了两个数量级的服务水平协议(SLA)违规。我们展示了如何使用针对不同SLA或特定应用程序训练的模型来优化PPW,例如,在现场改进数据中心硬件。由此产生的CPU第一次满足了真实的部署标准,并提供了一种新的方法来为单个客户定制硬件,即使他们的需求发生了变化。
Processors that adapt architecture to workloads at runtime promise compelling performance per watt (PPW) gains, offering one way to mitigate diminishing returns from pipeline scaling. State-of-the-art adaptive CPUs deploy machine learning (ML) models on-chip to optimize hardware by recognizing workload patterns in event counter data. However, despite breakthrough PPW gains, such designs are not yet widely adopted due to the potential for systematic adaptation errors in the field. This paper presents an adaptive CPU based on Intel SkyLake that (1) closes the loop to deployment, and (2) provides a novel mechanism for post-silicon customization. Our CPU performs predictive cluster gating, dynamically setting the issue width of a clustered architecture while clock-gating unused resources. Gating decisions are driven by ML adaptation models that execute on an existing microcontroller, minimizing design complexity and allowing performance characteristics to be adjusted with the ease of a firmware update. Crucially, we show that although adaptation models can suffer from statistical blindspots that risk degrading performance on new workloads, these can be reduced to minimal impact with careful design and training. Our adaptive CPU improves PPW by 31.4% over a comparable non-adaptive CPU on SPEC2017, and exhibits two orders of magnitude fewer Service Level Agreement (SLA) violations than the state-of-the-art. We show how to optimize PPW using models trained to different SLAs or to specific applications, e.g. to improve datacenter hardware in situ. The resulting CPU meets real world deployment criteria for the first time and provides a new means to tailor hardware to individual customers, even as their needs change.