Diagnosing Distributed Systems with Self-propelled Instrumentation

Diagnosing Distributed Systems with Self-propelled Instrumentation
复制标题

使用自走式仪器诊断分布式系统

DOI:
--
复制
发表时间:
2008
期刊:
International Middleware Conference
影响因子:
--
通讯作者:
B. Miller
B. Miller
中科院分区:
--
文献类型:
--
作者:
Alexander V. Mirgorodskiy;B. Miller

文献摘要

被引文献

相似文献

我们提出了一个由三部分组成的方法来诊断生产分布式环境中的错误和性能问题。首先,我们介绍了一种新的执行监控技术,动态注入的代码片段,代理,到一个应用程序的需求。代理在进程内的控制流之前插入检测,并传播到其他进程中,跟踪通信事件,跨越主机边界,并收集执行的分布式函数级跟踪。其次,我们提出了一个算法,分离到用户有意义的活动称为流的跟踪。此步骤简化了手动检查并实现了跟踪的自动分析。最后,我们描述了我们的自动化的根本原因分析技术,比较流,以帮助分析师找到一个异常的流,并确定该流中的功能,这是一个可能的原因异常。我们通过诊断Condor分布式调度系统中的两个复杂问题来证明我们技术的有效性。
We present a three-part approach for diagnosing bugs and performance problems in production distributed environments. First, we introduce a novel execution monitoring technique that dynamically injects a fragment of code, the agent, into an application process on demand. The agent inserts instrumentation ahead of the control flow within the process and propagates into other processes, following communication events, crossing host boundaries, and collecting a distributed function-level trace of the execution. Second, we present an algorithm that separates the trace into user-meaningful activities called flows. This step simplifies manual examination and enables automated analysis of the trace. Finally, we describe our automated root cause analysis technique that compares the flows to help the analyst locate an anomalous flow and identify a function in that flow that is a likely cause of the anomaly. We demonstrate the effectiveness of our techniques by diagnosing two complex problems in the Condor distributed scheduling system.