Log Discovery for Troubleshooting Open Distributed Systems with TLQ

Log Discovery for Troubleshooting Open Distributed Systems with TLQ
复制标题

DOI:
10.1145/3311790.3396633
复制
发表时间:
2020-07
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
Nathaniel Kremer-Herman;D. Thain
Nathaniel Kremer-Herman;D. Thain
中科院分区:
其他
文献类型:
--
作者:
Nathaniel Kremer-Herman;D. Thain

文献摘要

相似文献

对分布式系统进行故障排除可能非常困难。期望用户知道他们的系统和系统中使用的每台机器的环境配置之间的细粒度交互是不可行的。正因为如此,当一个看似微不足道的细节发生变化时,工作可能会陷入停顿。为了解决这个问题,有大量最先进的日志分析工具、调试器和可视化套件。然而,用户可以在开放分布式系统中执行,其中在运行时之前不知道其组件的放置。这使得跟踪调试日志的过程几乎与对这些日志记录的故障进行故障排除一样困难,因为这些日志的位置通常对用户(以及他们正在使用的故障排除工具)不透明。我们提出TLQ,日志发现的第一原则,使开放的分布式系统故障排除设计的框架。TLQ由一个查询客户端和一组服务器组成,这些服务器跟踪分布在开放分布式系统中的相关调试日志。通过一系列的例子,我们演示了如何TLQ使用户能够发现他们的系统的调试日志的位置,并反过来使用定义良好的故障排除工具,这些日志在一个分布式的方式。这两个任务以前是不切实际的要求一个开放的分布式系统没有显着的先验知识。我们还通过生产系统具体验证了TLQ的有效性:生物多样性科学工作流程。我们注意到潜在的存储和性能开销TLQ相比,一个集中的,封闭的系统方法。
Troubleshooting a distributed system can be incredibly difficult. It is rarely feasible to expect a user to know the fine-grained interactions between their system and the environment configuration of each machine used in the system. Because of this, work can grind to a halt when a seemingly trivial detail changes. To address this, there is a plethora of state-of-the-art log analysis tools, debuggers, and visualization suites. However, a user may be executing in an open distributed system where the placement of their components are not known before runtime. This makes the process of tracking debug logs almost as difficult as troubleshooting the failures these logs have recorded because the location of those logs is usually not transparent to the user (and by association the troubleshooting tools they are using). We present TLQ, a framework designed from first principles for log discovery to enable troubleshooting of open distributed systems. TLQconsists of a querying client and a set of servers which track relevant debug logs spread across an open distributed system. Through a series of examples, we demonstrate how TLQenables users to discover the locations of their system’s debug logs and in turn use well-defined troubleshooting tools upon those logs in a distributed fashion. Both of these tasks were previously impractical to ask of an open distributed system without significant a priori knowledge. We also concretely verify TLQ’s effectiveness by way of a production system: a biodiversity scientific workflow. We note the potential storage and performance overheads of TLQcompared to a centralized, closed system approach.