Guaranteeing fault tolerance through scheduling in real-time systems

Guaranteeing fault tolerance through scheduling in real-time systems
复制标题

DOI:
--
复制
发表时间:
1996
期刊:
--
影响因子:
--
通讯作者:
D. Mossé;Sunondo Ghosh
D. Mossé;Sunondo Ghosh
中科院分区:
其他
文献类型:
--
作者:
D. Mossé;Sunondo Ghosh

文献摘要

被引文献

相似文献

实时系统是那些必须在其时间限制内执行所有任务的系统。由于一些实时任务错过最后期限的灾难性后果,容错是此类系统的重要组成部分。本文介绍了通过引入时间冗余来提高实时系统容错能力的技术。时间冗余在超可靠实时系统中是必不可少的,在超可靠实时系统中,相关故障必须被容忍。它也可以用来检测和容忍瞬态故障,这是大多数的计算系统中的故障。本论文展示了如何将时间冗余与硬件和软件冗余结合起来,以容忍实时系统中的各种故障。本文考虑了几个不同的系统和任务模型,并为每个模型,提出了一个可扩展性测试(一个利用界或一组条件),保证系统中的所有任务将满足其时间约束,即使在故障的存在。本文研究了系统的容错能力和资源利用率(由于增加冗余而降低)之间的权衡。新技术的引入,以提高系统的利用率。给出了有效的调度算法和调度界限,保证了任务的高可调度性。对本文提出的容错方法进行了全面的评价。对于静态和动态系统,测量系统从一个故障恢复并准备容忍第二个故障的时间。进行各种tradeo研究,以帮助系统设计人员做出适当的选择。大量的仿真结果解释了各种输入参数,如任务特性,
Real-time systems are those which must execute all tasks within their timing constraints. Due to the catastrophic consequences of missing deadlines of some realtime tasks, fault tolerance is an essential component of such systems. This thesis introduces techniques to enhance the fault tolerance capability of real-time systems by incorporating time redundancy. Time redundancy is essential in ultrareliable real-time systems where correlated faults must be tolerated. It can also be used to detect and tolerate transient faults, which are a majority of the faults in computing systems. This thesis demonstrates how time redundancy can be used in conjunction with hardware and software redundancy to tolerate a variety of faults in real-time systems. This thesis considers several di erent system and task models, and for each model, presents a schedulability test (a utilization bound or a set of conditions) which guarantees that all tasks in the system will satisfy their timing constraints even in the presence of faults. The thesis studies the tradeo between the fault tolerance capability and resource utilization of the system (which decreases due to the added redundancy). New techniques are introduced to increase the system utilization. Efcient scheduling algorithms and bounds are presented to ensure high schedulability of tasks. The fault tolerance approaches presented in this thesis are thoroughly evaluated. The time after which a system recovers from one fault and is ready to tolerate a second one is measured for static and dynamic systems. Various tradeo studies are conducted to help the system designer make appropriate choices. Extensive simulation results explain the e ects of various input parameters such as task characteristics,