Localizing Runtime Anomalies in Service-Oriented Systems

Localizing Runtime Anomalies in Service-Oriented Systems
复制标题

DOI:
10.1109/tsc.2016.2593462
复制
发表时间:
2017-01-01
影响因子:
8.1
通讯作者:
Yang, Yun
Yang, Yun
中科院分区:
计算机科学2区
文献类型:
--
作者:
He, Qiang;Xie, Xiaoyuan;Yang, Yun

文献摘要

被引文献

相似文献

在分布式、动态和易变的操作环境中,面向服务的系统(SoSS)中发生的运行时异常必须被及时定位和修复,以保证响应用户请求的结果的成功交付。持续监视所有组件服务并检查整个SOS的运行时异常是不切实际的,因为需要过多的资源和时间消耗,尤其是在大规模场景中。我们提出了一种基于频谱的方法,该方法经历了五个阶段的过程,以基于端到端系统延迟快速定位SOSS中发生的运行时异常。对于运行时异常,我们的方法计算SOS的每个基本组件(BC)的相似系数来评估它们是否有故障。我们的方法还计算延迟系数来评估每个BC对端到端系统延迟严重程度的贡献。最后,根据BCS的相似系数得分和延迟系数得分进行排序,以确定它们被检查的顺序。通过大量实验验证了该方法的有效性和高效性。结果表明,我们的方法在定位单个和多个运行时异常方面明显优于随机检查和流行的基于Ochiai的检查。因此,我们的方法可以帮助节省时间和精力来定位在SoS中发生的运行时异常。
In a distributed, dynamic and volatile operating environment, runtime anomalies occurring in service-oriented systems (SOSs) must be located and fixed in a timely manner in order to guarantee successful delivery of outcomes in response to user requests. Monitoring all component services constantly and inspecting the entire SOS upon a runtime anomaly are impractical due to excessive resource and time consumption required, especially in large-scale scenarios. We present a spectrum-based approach that goes through a five-phase process to quickly localize runtime anomalies occurring in SOSs based on end-to-end system delays. Upon runtime anomalies, our approach calculates the similarity coefficient for each basic component (BC) of the SOS to evaluate their suspiciousness of being faulty. Our approach also calculates the delay coefficients to evaluate each BC's contribution to the severity of the end-to-end system delays. Finally, the BCs are ranked by their similarity coefficient scores and delay coefficient scores to determine the order of them being inspected. Extensive experiments are conducted to evaluate the effectiveness and efficiency of the proposed approach. The results indicate that our approach significantly outperforms random inspection and the popular Ochiai-based inspection in localizing single and multiple runtime anomalies effectively. Thus, our approach can help save time and effort for localizing runtime anomalies occuring in SOSs.