Self-healing multitier architectures using cascading rescue points

Self-healing multitier architectures using cascading rescue points
复制标题

使用级联救援点的自我修复多层架构

DOI:
--
复制
发表时间:
2012
期刊:
Asia-Pacific Computer Systems Architecture Conference
影响因子:
--
通讯作者:
A. Keromytis
A. Keromytis
中科院分区:
--
文献类型:
--
作者:
Angeliki Zavou;G. Portokalidis;A. Keromytis

文献摘要

被引文献

相似文献

软件缺陷和漏洞会给家庭用户和互联网基础设施带来严重问题,限制互联网服务的可用性,导致数据丢失,并降低系统完整性。使用救援点(RPS)的软件自我修复是一种从不可预见的错误中恢复的已知机制。然而,在多层体系结构上应用它可能会有问题,因为某些操作,如通过网络传输数据,无法撤消。我们提出级联救援点(CRP)来解决在互连应用中使用传统救援点从错误中恢复时可能出现的状态不一致问题。使用CRPS,当在RP内执行的应用程序传输数据时,远程对等体被通知也执行检查点,因此通信实体以协调但松散耦合的方式进行检查点。当RPS成功完成执行和启动恢复时,也会发送通知,以便远程方执行适当的操作。我们开发了一个工具,通过动态检测二进制文件并在应用程序之间已建立的TCP通道中透明地插入通知来实现CRPS。我们在不同的应用程序上测试了我们的工具,包括MySQL和Apache服务器,结果表明,它允许它们成功地从错误中恢复,同时产生的开销在4.54%到71.56%之间。
Software bugs and vulnerabilities cause serious problems to both home users and the Internet infrastructure, limiting the availability of Internet services, causing loss of data, and reducing system integrity. Software self-healing using rescue points (RPs) is a known mechanism for recovering from unforeseen errors. However, applying it on multitier architectures can be problematic because certain actions, like transmitting data over the network, cannot be undone. We propose cascading rescue points (CRPs) to address the state inconsistency issues that can arise when using traditional RPs to recover from errors in interconnected applications. With CRPs, when an application executing within a RP transmits data, the remote peer is notified to also perform a checkpoint, so the communicating entities checkpoint in a coordinated, but loosely coupled way. Notifications are also sent when RPs successfully complete execution, and when recovery is initiated, so that the appropriate action is performed by remote parties. We developed a tool that implements CRPs by dynamically instrumenting binaries and transparently injecting notifications in the already established TCP channels between applications. We tested our tool with various applications, including the MySQL and Apache servers, and show that it allows them to successfully recover from errors, while incurring moderate overhead between 4.54% and 71.56%.