Recovery domains: an organizing principle for recoverable operating systems

Recovery domains: an organizing principle for recoverable operating systems
复制标题

恢复域:可恢复操作系统的组织原则

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Samuel T. King
Samuel T. King
中科院分区:
--
文献类型:
--
作者:
Andrew Lenharth;Vikram S. Adve;Samuel T. King

文献摘要

被引文献

相似文献

我们描述了一种策略,使现有的商用操作系统能够从内核几乎任何部分(包括核心内核组件)的意外运行时错误中恢复过来。我们的方法是动态的和面向请求的;它将故障的影响隔离到导致故障的请求,而不是隔离到静态内核组件。这种方法基于“恢复域”的概念,这是一种组织原则,可以在多线程系统中对受请求影响的状态进行回滚,同时对其他请求或线程的影响最小。我们在Linux内核的v2.4.22和v2.6.27上应用了这种方法,它需要修改或新建132行代码:其他更改都是通过编译器的简单插装传递来执行的。我们的实验表明,该方法能够在恢复事件期间以最小的附带影响从其他致命故障中恢复。
We describe a strategy for enabling existing commodity operating systems to recover from unexpected run-time errors in nearly any part of the kernel, including core kernel components. Our approach is dynamic and request-oriented; it isolates the effects of a fault to the requests that caused the fault rather than to static kernel components. This approach is based on a notion of "recovery domains," an organizing principle to enable rollback of state affected by a request in a multithreaded system with minimal impact on other requests or threads. We have applied this approach on v2.4.22 and v2.6.27 of the Linux kernel and it required 132 lines of changed or new code: the other changes are all performed by a simple instrumentation pass of a compiler. Our experiments show that the approach is able to recover from otherwise fatal faults with minimal collateral impact during a recovery event.