Definition and specification of accrual failure detectors

Definition and specification of accrual failure detectors
复制标题

DOI:
10.1109/dsn.2005.37
复制
发表时间:
2005-03
期刊:
2005 International Conference on Dependable Systems and Networks (DSN'05)
影响因子:
--
通讯作者:
X. Défago;P. Urbán;Naohiro Hayashibara;T. Katayama
X. Défago;P. Urbán;Naohiro Hayashibara;T. Katayama
中科院分区:
其他
文献类型:
--
作者:
X. Défago;P. Urbán;Naohiro Hayashibara;T. Katayama

文献摘要

被引文献

相似文献

多年来,人们一直倡导将故障检测作为一项基本服务来发展,但遗憾的是,迄今为止并没有取得多大的成功。我们认为,这是因为重要的系统工程问题尚未得到充分解决,从而阻碍了真正通用服务的定义。最终,我们的目标是定义一个既简单又富有表现力的服务,但又足够强大以支持许多分布式应用程序的要求。为此,我们考虑一种替代的服务和应用程序之间的交互模型,称为应计故障检测器。粗略地说,应计故障检测器将表示怀疑级别的真实的值而不是传统的二进制信息(即,信任vs.怀疑)。在本文中,我们提供了一个严格的定义应计故障检测器,表明改变交互模型导致计算能力没有损失,讨论服务质量问题,并提出了几种可能的实现。
For many years, people have been advocating the development of failure detection as a basic service, but, unfortunately, without meeting much success so far. We believe that this comes from the fact that important system engineering issues have not yet been addressed adequately, thus preventing the definition of a truly generic service. Ultimately, our goal is to define a service that is both simple and expressive, yet powerful enough to support the requirements of many distributed applications. To this end, we consider an alternative interaction model between the service and the applications, called accrual failure detectors. Roughly, an accrual failure detector associates to each process a real value representing a suspicion level, instead of the traditional binary information (i.e., trust vs. suspect). In this paper, we provide a rigorous definition for accrual failure detectors, demonstrate that changing the interaction model leads to no loss in computational power, discuss quality of service issues, and present several possible implementations.