Towards Pre-Deployment Detection of Performance Failures in Cloud Distributed Systems

Towards Pre-Deployment Detection of Performance Failures in Cloud Distributed Systems
复制标题

云分布式系统性能故障的部署前检测

DOI:
--
复制
发表时间:
2015
期刊:
USENIX Workshop on Hot Topics in Cloud Computing
影响因子:
--
通讯作者:
Haryadi S. Gunawi
Haryadi S. Gunawi
中科院分区:
--
文献类型:
--
作者:
Riza O. Suminto;Agung Laksono;A. Satria;Thanh Do;Haryadi S. Gunawi

文献摘要

被引文献

相似文献

现代分布式系统(“云系统”)已成为当今许多应用程序的主导支柱。它们有不同的形式,例如横向扩展文件系统、键值存储、计算框架、同步和集群管理服务。随着这些系统共同成为“云操作系统”,用户期望包括性能稳定性在内的高可靠性。不幸的是,它们必须运行的软件和环境的复杂性已经超过了现有的测试和调试工具。云系统必须以不同的拓扑规模运行,执行复杂的分布式协议,面对负载波动和广泛的硬件故障,并为具有不同工作特征的用户提供服务。一类重要的故障是性能故障,即系统(例如 Hadoop)无法提供预期性能的情况(例如,作业花费的时间比平时长 10 倍)。与云工程师的对话反映,性能稳定性往往比性能优化更重要;当发生性能故障时,用户会感到沮丧,系统会浪费和未充分利用资源,并且需要长时间的调试工作才能发现并解决问题。遗憾的是,性能故障仍然很常见。我们之前的工作表明,云系统开发人员报告的重要问题中有 22% 与性能错误有关 [12]。在本文中,我们的重点是回答以下三个问题:云系统中出现的性能错误的根本原因是什么?目前检测性能缺陷的技术还缺少什么?有哪些新的方向可以防止现场发生性能故障?
Modern distributed systems (“cloud systems”) have emerged as a dominant backbone for many today’s applications. They come in different forms such as scale-out file systems, key-value stores, computing frameworks, synchronization and cluster management services. As these systems collectively become the “cloud operating system”, users expect high dependability including performance stability. Unfortunately, the complexity of the software and environment in which they must run has outpaced existing testing and debugging tools. Cloud systems must run at scale with different topologies, execute complex distributed protocols, face load fluctuations and a wide range of hardware faults, and serve users with diverse job characteristics. One type of important failures is performance failures, a situation where a system (e.g., Hadoop) does not deliver the expected performance (e.g., a job takes 10x longer time than usual). Conversation with cloud engineers reflects that performance stability is often more important than performance optimization; when performance failures happen, users are frustrated, systems waste and underutilize resources, and long debugging efforts are required to find and fix the problems. Sadly, performance failures are still common; our previous work shows that 22% of vital issues reported by cloud system developers relate to performance bugs [12]. In this paper, our focus is to answer the following three questions: What is the root-cause anatomy of performance bugs that appear in cloud systems? What is missing within the state of the art of detecting performance bugs? What are new novel directions that can prevent performance failures to happen in the field?