A comprehensive model for software rejuvenation

A comprehensive model for software rejuvenation
复制标题

DOI:
10.1109/tdsc.2005.15
复制
发表时间:
2005-04
影响因子:
7.3
通讯作者:
K. Vaidyanathan;Kishor S. Trivedi
K. Vaidyanathan;Kishor S. Trivedi
中科院分区:
计算机科学2区
文献类型:
--
作者:
K. Vaidyanathan;Kishor S. Trivedi

文献摘要

被引文献

相似文献

近来,已经报道了软件老化现象,其中软件系统的状态随着时间而劣化。这种现象可能最终导致系统性能下降和/或崩溃/挂起故障,是操作系统资源耗尽、数据损坏和数值错误累积的结果。为了对抗软件老化,已经提出了一种称为软件恢复的技术,其本质上涉及偶尔终止应用程序或系统,清理其内部状态和/或其环境,并重新启动它。由于恢复会产生开销,因此一个重要的研究问题是确定启动此操作的最佳时间。在本文中,我们首先描述了如何将故障归因于软件老化的框架中的灰色的软件故障分类(确定性和瞬态),并研究的治疗和恢复策略,为每个故障类。然后,我们构建了一个半马尔可夫奖励模型的基础上收集的UNIX操作系统的工作负载和资源使用数据。我们确定不同的工作负载状态,使用统计聚类分析,估计转移概率,和逗留时间分布的数据。对应于每个资源,然后基于每个状态中的资源消耗率为模型定义奖励函数。然后求解该模型以获得每个资源的估计耗尽时间。然后,半马尔可夫奖励模型的结果被馈送到更高级别的可用性模型中,该模型考虑了故障,然后是被动恢复和主动恢复。这个全面的模型,然后用来获得最佳的复兴时间表,最大限度地提高可用性或最小化停机成本。
Recently, the phenomenon of software aging, one in which the state of the software system degrades with time, has been reported. This phenomenon, which may eventually lead to system performance degradation and/or crash/hang failure, is the result of exhaustion of operating system resources, data corruption, and numerical error accumulation. To counteract software aging, a technique called software rejuvenation has been proposed, which essentially involves occasionally terminating an application or a system, cleaning its internal state and/or its environment, and restarting it. Since rejuvenation incurs an overhead, an important research issue is to determine optimal times to initiate this action. In this paper, we first describe how to include faults attributed to software aging in the framework of Gray's software fault classification (deterministic and transient), and study the treatment and recovery strategies for each of the fault classes. We then construct a semi-Markov reward model based on workload and resource usage data collected from the UNIX operating system. We identify different workload states using statistical cluster analysis, estimate transition probabilities, and sojourn time distributions from the data. Corresponding to each resource, a reward function is then defined for the model based on the rate of resource depletion in each state. The model is then solved to obtain estimated times to exhaustion for each resource. The result from the semi-Markov reward model are then fed into a higher-level availability model that accounts for failure followed by reactive recovery, as well as proactive recovery. This comprehensive model is then used to derive optimal rejuvenation schedules that maximize availability or minimize downtime cost.