Reliability-Aware Runtime Adaption Through a Statically Generated Task Schedule

Reliability-Aware Runtime Adaption Through a Statically Generated Task Schedule
复制标题

DOI:
10.1109/tvlsi.2017.2753242
复制
发表时间:
2018
影响因子:
2.8
通讯作者:
Laura Rozo;A. Landwehr;Yan Zheng;Chengmo Yang;Guangrong Gao
Laura Rozo;A. Landwehr;Yan Zheng;Chengmo Yang;Guangrong Gao
中科院分区:
工程技术2区
文献类型:
--
作者:
Laura Rozo;A. Landwehr;Yan Zheng;Chengmo Yang;Guangrong Gao

文献摘要

被引文献

相似文献

器件的尺寸缩小、单个芯片中组件数量的增加、环境问题的变化以及老化效应带来了严重的可靠性挑战,这些挑战对系统的操作施加了严格的约束。为了科普这些挑战,本文提出了一个可靠性感知调度框架,该框架结合了静态和动态分析,以提高整个系统对不同类型故障的弹性(即,间歇的、瞬时的和永久的)。静态分析技术采用遗传算法来优化整个系统的可靠性,考虑可靠性水平(RL)作为一个中间调度维度,并创建一个任务到RL映射。这使得RL到核心的映射能够在运行时根据故障率变化进行有效地调整,而任务到RL的映射仍然可以被重用。动态分析跟踪出现在每个核心中的故障,并测量这些故障的时间相关性,以更新RL到核心的映射。所提出的可靠性感知框架实现在一个国家的最先进的运行时系统,特拉华州自适应运行时系统,从而定量地显示在现有的多核平台上使用的整体框架的优势。实验结果表明,所提出的技术可将应用程序执行时间提高高达30%,并将运行时发生的错误提高高达72%。
Device scaling, increasing number of components in a single chip, varying environmental issues, and aging effects have brought severe reliability challenges that impose tight constraints on the operation of a system. To cope with these challenges, this paper proposes a reliability-aware scheduling framework that combines static and dynamic analyses to improve the overall system resiliency to different kinds of faults (i.e., intermittent, transient, and permanent). The static analysis technique employs genetic algorithms to optimize the overall system reliability by considering reliability level (RL) as an intermediate scheduling dimension and creating a task-to-RL mapping. This enables the RL-to-core mapping to be efficiently adapted at runtime according to fault rate variations, while the task-to-RL mapping can still be reused. The dynamic analysis tracks faults appearing in each core and measures the time correlation of those faults to update the RL-to-core mapping. The proposed reliability-aware framework is implemented in a state-of-the-art runtime system, Delaware Adaptive Run-Time System, so as to quantitatively show the advantages of using the overall framework in existing multicore platforms. Experimental results show that the proposed technique delivers up to 30% improvement in application execution time and up to 72% improvement in faults occurring at runtime.