Energy Efficient Lifetime Reliability-Aware Checkpointing for Real-Time System

Energy Efficient Lifetime Reliability-Aware Checkpointing for Real-Time System
复制标题

实时系统的节能终生可靠性感知检查点

DOI:
10.1166/jolpe.2014.1343
复制
发表时间:
2014
影响因子:
--
通讯作者:
Bandan M
Bandan M
中科院分区:
--
文献类型:
--
作者:
Bandan M

文献摘要

被引文献

相似文献

由于技术的不断扩展,集成电路(IC)的可靠性是一个新兴的设计挑战,特别是在各种不同的工作环境中。现代系统的寿命可靠性受到较高的磨损和应力效应的严重限制。检查点作为一种有效的容错系统设计方法,已经得到了广泛的应用。传统上,它通过在预定义的时间保存中间结果并在需要时回滚到适当的先前保存的状态来容忍暂态故障的影响。在本文中,我们提出了一种新的双工实时系统的检查点机制,实现了对瞬时和永久故障的容错,并提供了一种通过将任务从不健康(可能接近死亡)的主机迁移到备用主机的故障避免机制。我们开发了一个数学模型来评估所提出的方法在各种故障和任务迁移情况下的性能。检查点和任务迁移相结合,通过容忍故障和损耗,提高了系统的生命周期可靠性。由于检查点会增加额外的开销,因此对于任何实时系统来说,能量消耗和满足任务截止日期的能力都是非常重要的。任务的预期执行时间(EET)是任务完成的重要性能指标。类似地,平均能耗(AEC)反映了检查点机制在各种故障下的能源使用情况。在各种故障的概率分布下,我们对所提出的检查点机制进行了EET和AEC的评估。我们还研究了我们提出的算法的最后期限估计。结果表明,在故障率高达10-3的情况下,该算法仍能满足最后期限的要求。仿真结果表明,所提出的检查点机制能够满足任务期限,时间开销仅为12.57%。
Due to continued technology scaling, reliability of today's integrated circuits (IC) is an emerging design challenge especially in varied range of operating environment. The lifetime reliability of modern system has been severely limited by higher wear-out and stress effects. Checkpointing has been extensively used as an effective method in fault-tolerant system design. Traditionally, it is used to tolerate the impact of transient faults through saving the intermediate results at predefined time and rolling-back to appropriate previously saved state whenever needed. In this paper, we proposed a new checkpointing mechanism for a duplex real-time system that achieves fault-tolerant against transient and permanent faults, and also provides a fault avoidance mechanism by migrating task from an unhealthy (perhaps near-to-die) host to a spare host. We developed a mathematical model for evaluating the performance of the proposed methodology in presence of various faults and task migration. The combination of checkpointing and task migration enhances the lifetime reliability of the system by tolerating faults and wear-out. Since checkpointing imposes additional overhead, energy consumption and ability to meet the task deadline are very crucial for any real-time system. The Expected-Execution-Time (EET) of a task is an important performance metric in respect to task completion. Similarly, the Average-Energy-Consumption (AEC) reflects the energy usage of a checkpointing mechanism under various faults. Under probabilistic distribution of various faults, we evaluate EET and AEC for our proposed checkpointing mechanism. We also investigated the deadline estimation for our proposed algorithm. We found that the proposed algorithm is able to meet the deadline even when the fault rate is as high as 10–3. Our simulation result shows that the proposed checkpointing mechanism can meet task deadline with only 12.57% time overhead.