Timely Long Tail Identification through Agent Based Monitoring and Analytics

Timely Long Tail Identification through Agent Based Monitoring and Analytics
复制标题

DOI:
10.1109/isorc.2015.39
复制
发表时间:
2015-04
期刊:
2015 IEEE 18th International Symposium on Real-Time Distributed Computing
影响因子:
--
通讯作者:
Peter Garraghan;Ouyang Xue;P. Townend;Jie Xu
Peter Garraghan;Ouyang Xue;P. Townend;Jie Xu
中科院分区:
其他
文献类型:
--
作者:
Peter Garraghan;Ouyang Xue;P. Townend;Jie Xu

文献摘要

被引文献

相似文献

分布式系统的复杂性和规模的不断增加导致了涌现行为的出现,这极大地影响了整个系统的性能。一个重要的紧急属性是“长尾”,即一小部分任务掉队者显着影响作业执行完成时间。为了减轻这种行为,需要及时准确地识别系统内发生的离散任务。然而,目前的方法侧重于缓解而不是识别,这通常在执行生命周期中识别掉队者太晚。本文提出了一种方法和工具,通过在线和离线分析相结合,及时识别分布式系统中的长尾行为。这是通过历史分析来分析和建模任务执行模式,然后通知在运行时监视任务执行的在线分析代理来实现的。此外,我们还对两个大规模生产云数据输入进行了实证分析,证明了现代分布式系统中数据偏斜的挑战,该分析表明,由数据偏斜引起的约5%的任务掉队影响了批处理总作业的50%。我们的研究结果表明,我们的方法能够以98%的准确率识别出不到11%的执行生命周期的任务掉队者,这意味着与当前最先进的实践相比有了显着的改进,并在全球范围内的大型分布式系统中实现了更有效的缓解策略。
The increasing complexity and scale of distributed systems has resulted in the manifestation of emergent behavior which substantially affects overall system performance. A significant emergent property is that of the "Long Tail", whereby a small proportion of task stragglers significantly impact job execution completion times. To mitigate such behavior, straggling tasks occurring within the system need to be accurately identified in a timely manner. However, current approaches focus on mitigation rather than identification, which typically identify stragglers too late in the execution lifecycle. This paper presents a method and tool to identify Long Tail behavior within distributed systems in a timely manner, through a combination of online and offline analytics. This is achieved through historical analysis to profile and model task execution patterns, which then inform online analytic agents that monitor task execution at runtime. Furthermore, we provide an empirical analysis of two large-scale production Cloud data enters that demonstrate the challenge of data skew within modern distributed systems, this analysis shows that approximately 5% of task stragglers caused by data skew impact 50% of the total jobs for batch processes. Our results demonstrate that our approach is capable of identifying task stragglers less than 11% into their execution lifecycle with 98% accuracy, signifying significant improvement over current state-of-the-art practice and enables far more effective mitigation strategies in large-scale distributed systems worldwide.