Failure Resilience for Device Drivers

Failure Resilience for Device Drivers
复制标题

设备驱动程序的故障恢复能力

DOI:
10.1109/dsn.2007.46
复制
发表时间:
2007
期刊:
Dependable Systems and Networks
影响因子:
--
通讯作者:
A. Tanenbaum
A. Tanenbaum
中科院分区:
--
文献类型:
--
作者:
J. Herder;H. Bos;Ben Gras;P. Homburg;A. Tanenbaum

文献摘要

被引文献

相似文献

研究表明,设备驱动程序和扩展程序包含的错误是其他操作系统代码的3-7倍,因此更容易失败。因此,我们提出了一个故障恢复操作系统的设计,可以从死驱动程序和其他关键组件恢复-主要是通过监控和更换故障组件的飞行-透明的应用程序,而无需用户干预。本文重点研究了尸体的复原程序。我们解释了我们的缺陷检测机制,政策驱动的恢复过程,并重新启动后的组件重新整合的工作。此外,我们还讨论了从网络、块设备和字符设备驱动程序故障中恢复的具体步骤。最后,我们评估我们的设计使用性能测量,软件故障注入实验,并分析了再工程的努力。
Studies have shown that device drivers and extensions contain 3-7 times more bugs than other operating system code and thus are more likely to fail. Therefore, we present a failure-resilient operating system design that can recover from dead drivers and other critical components - primarily through monitoring and replacing malfunctioning components on the fly - transparent to applications and without user intervention. This paper focuses on the post-mortem recovery procedure. We explain the working of our defect detection mechanism, the policy-driven recovery procedure, and post-restart reintegration of the components. Furthermore, we discuss the concrete steps taken to recover from network, block device, and character device driver failures. Finally, we evaluate our design using performance measurements, software fault-injection experiments, and an analysis of the reengineering effort.