Profiling distributed systems in lightweight virtualized environments with logs and resource metrics

Profiling distributed systems in lightweight virtualized environments with logs and resource metrics
复制标题

使用日志和资源指标分析轻量级虚拟化环境中的分布式系统

DOI:
10.1145/3208040.3208044
复制
发表时间:
2018
期刊:
Proceedings of the 27th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
Mike Ji
Mike Ji
中科院分区:
--
文献类型:
--
作者:
Aidi Pi;Wei Chen;Xiaobo Zhou;Mike Ji

文献摘要

被引文献

相似文献

理解和故障排除云中的分布式系统被认为是一个非常困难的问题,因为单个用户请求的执行分配给了多台计算机。此外,云环境的多租期性质进一步引入了导致性能问题的干扰。大多数现有的故障排除工具要么专注于日志分析或侵入性跟踪方法,因此未探索资源使用率监视。我们提出并实施LRTRACE,这是一种非侵入性的跟踪和反馈控制工具,用于轻巧虚拟化环境中的分布式应用程序。 LRTRACE以细粒度的方式在运行时介绍了应用程序的日志消息和实际资源消耗,这是通过基于容器的轻质容器虚拟化使其成为可能的。通过将这两种信息关联,LRTRACE为用户提供了建立资源消费和应用程序事件变化之间关系的能力。此外,LRTRACE允许用户定义和实现自己的反馈控制插件,以半自动的方式管理群集。在系统评估中,我们在多租户集群中运行Spark和MapReduce应用程序,并表明LRTRACE可以诊断由干扰或错误或两者兼而有之引起的性能问题。它还可以帮助用户了解数据并行应用程序的工作流程。
Understanding and troubleshooting distributed systems in the cloud is considered a very difficult problem because the execution of a single user request is distributed to multiple machines. Further, the multi-tenancy nature of cloud environments further introduces interference that causes performance issues. Most existing troubleshooting tools either focus on log analysis or intrusive tracing methods, leaving resource usage monitoring unexplored. We propose and implement LRTrace, a non-intrusive tracing and feedback control tool for distributed applications in lightweight virtualized environments. LRTrace profiles both log messages and actual resource consumptions of an application at runtime in a fine-grained manner, which is made possible by lightweight container-based virtualization. By correlating these two kinds of information, LRTrace provides users the ability to build the relationship between changes in resource consumption and application events. Furthermore, LRTrace allows users to define and implement their own feedback control plug-ins to manage the cluster in a semi-automatic manner. In system evaluation, we run Spark and MapReduce applications in a multi-tenant cluster and show that LRTrace can diagnose performance issues caused by either interference or bugs, or both. It also helps users to understand the workflows of data-parallel applications.