WARM: Workload-Aware Reliability Management in Linux/Android

WARM: Workload-Aware Reliability Management in Linux/Android
复制标题

WARM:Linux/Android 中的工作负载感知可靠性管理

DOI:
10.1109/tcad.2016.2611501
复制
发表时间:
2017
影响因子:
2.9
通讯作者:
T. Rosing
T. Rosing
中科院分区:
计算机科学3区
文献类型:
--
作者:
Pietro Mercati;Francesco Paterna;Andrea Bartolini;L. Benini;T. Rosing

文献摘要

被引文献

相似文献

随着CMOS尺寸超过14 nm,可靠性成为IC制造商的主要关注点。可靠性感知设计具有不可忽略的开销,并且不能考虑移动的设备中的用户体验。另一种方法是动态可靠性管理(DRM),它通过在运行时调整操作条件来抵消降级。在本文中,我们第一次制定DRM作为一个优化问题,占可靠性,温度和性能。我们开发了一个最佳的多核策略,使用凸优化,并表明它是不可行的,以实现在真实的系统。为此,我们提出了工作负载感知的可靠性管理(WARM),一种快速的DRM技术,适应不同的工作负载要求,以贸易的可靠性和用户体验。WARM在真实的Android设备上实现并测试。WARM平均在5%以内近似凸解算器的解,同时执行速度快了400多美元。WARM集成了一个热控制器,可以分配任务以满足热约束。这是必需的,因为降解强烈依赖于温度。我们表明,WARM满足温度限制在5%以内的情况下,比最先进的87.5%。我们表明,WARM任务分配实现了一年的寿命延长多核平台。它可以在集群架构(如big.LITTLE)上实现高达100%的性能改进,同时仍然保证可靠性目标。最后,我们表明,它实现了性能的4%的最大范围的应用,同时满足可靠性的限制。
With CMOS scaling beyond 14 nm, reliability is a major concern for IC manufacturers. Reliability-aware design has a non-negligible overhead and cannot account for user experience in mobile devices. An alternative is dynamic reliability management (DRM), which counteracts degradation by adapting the operating conditions at runtime. In this paper, for the first time we formulate DRM as an optimization problem that accounts for reliability, temperature and performance. We develop an optimal policy for multicores using convex optimization, and show that it is not feasible to implement on real systems. For this reason, we propose workload-aware reliability management (WARM), a fast DRM technique adapting to diverse workload requirements to trade reliability and user experience. WARM is implemented and tested on a real Android device. WARM approximates the solution of the convex solver within 5% on average, while executing more than $400 {\times }$ faster. WARM integrates a thermal controller that allocates tasks to meet thermal constraints. This is required since degradation strongly depends on temperature. We show that WARM meets temperature constraints within 5% in 87.5% more cases than the state-of-the-art. We show that WARM task allocation achieves up to one year lifetime improvement for a multicore platform. It can achieve up to 100% of performance improvement on cluster architectures, such as big.LITTLE, while still guaranteeing the reliability target. Finally, we show that it achieves performance in the 4% of the maximum for a broad range of a applications, while meeting the reliability constraints.