SAFER: System-level Architecture for Failure Evasion in Real-time Applications

SAFER: System-level Architecture for Failure Evasion in Real-time Applications
复制标题

SAFER:实时应用程序中故障规避的系统级架构

DOI:
--
复制
发表时间:
2012
期刊:
IEEE Real-Time Systems Symposium
影响因子:
--
通讯作者:
M. Jochim
M. Jochim
中科院分区:
--
文献类型:
--
作者:
Junsung Kim;Gaurav Bhatia;R. Rajkumar;M. Jochim

文献摘要

被引文献

相似文献

分布式嵌入式实时系统日益复杂的最新趋势给设计和实现可靠的系统(如自动驾驶汽车)带来了挑战。提高可靠性的传统方法是使用冗余硬件来复制整个(子)系统。尽管硬件复制已经广泛部署在硬实时系统中,如航空电子设备、航天飞机和核电站,但它对许多应用的吸引力明显降低,因为必要的硬件数量会随着系统规模的增加而增加。灵活系统设计的日益增长的需求也与硬件复制技术不一致。为了通过冗余实时操作来满足可靠性的需求,我们提出了一个称为SAFER(实时应用程序中故障规避的系统级架构)的层,该层结合了可配置的任务级容错功能,以容忍分布式嵌入式实时系统的故障停止处理器和任务故障。为了检测此类故障,SAFER监视每个任务的健康状态和状态信息,并广播这些信息。当使用基于时间的故障检测或基于事件的故障检测检测到故障时,SAFER会重新配置系统,以保留整个系统的功能。我们提供了更安全特性的最坏时序行为的形式化分析。我们还描述了一个配备了SAFER的系统的建模,通过基于模型的设计工具SysWeaver来分析时序特性。SAFER已在Ubuntu 10.04 LTS上实现,并部署在Boss上,Boss是卡内基梅隆大学开发的获奖自动驾驶汽车。我们展示了2007年DARPA城市挑战赛期间使用的各种模拟场景的测量结果。最后,我们给出了一个案例研究,当节点故障被注入时,更安全的故障恢复。
Recent trends towards increasing complexity in distributed embedded real-time systems pose challenges in designing and implementing a reliable system such as a self-driving car. The conventional way of improving reliability is to use redundant hardware to replicate the whole (sub)system. Although hardware replication has been widely deployed in hard real-time systems such as avionics, space shuttles and nuclear power plants, it is significantly less attractive to many applications because the amount of necessary hardware multiplies as the size of the system increases. The growing needs of flexible system design are also not consistent with hardware replication techniques. To address the needs of dependability through redundancy operating in real-time, we propose a layer called SAFER(System-level Architecture for Failure Evasion in Real-time applications) to incorporate configurable task-level fault-tolerance features to tolerate fail-stop processor and task failures for distributed embedded real-time systems. To detect such failures, SAFER monitors the health status and state information of each task and broadcasts the information. When a failure is detected using either time-based failure detection or event-based failure detection, SAFER reconfigures the system to retain the functionality of the whole system. We provide a formal analysis of the worst-case timing behaviors of SAFER features. We also describe the modeling of a system equipped with SAFER to analyze timing characteristics through a model-based design tool called SysWeaver. SAFER has been implemented on Ubuntu 10.04 LTS and deployed on Boss, an award-winning autonomous vehicle developed at Carnegie Mellon University. We show various measurements using simulation scenarios used during the 2007 DARPA Urban Challenge. Finally, we present a case study of failure recovery by SAFER when node failures are injected.