Straggler Mitigation at Scale

Straggler Mitigation at Scale
复制标题

DOI:
10.1109/tnet.2019.2946464
复制
发表时间:
2019-12-01
影响因子:
3.7
通讯作者:
Soljanin, Emina
Soljanin, Emina
中科院分区:
计算机科学2区
文献类型:
--
作者:
Aktas, Mehmet Fatih;Soljanin, Emina

文献摘要

被引文献

相似文献

分布式系统的性能可变性一直是一个主要问题,阻碍了现代分布式系统的可预测性和可扩展性。在多个服务器上冗余地执行请求或作业已被证明在理论和实践上都能有效地减轻可变性。采用冗余的系统引起了极大的关注,许多论文分析了在各种服务模型和运行时可变性假设下冗余的痛苦和收益。本文提出了一个成本(痛苦)与延迟(增益)的分析,执行多个任务的作业,采用复制或擦除编码冗余。服务时间的变化性的尾部沉重的痛苦和冗余的收益是决定性的,我们量化其影响,推导出的成本和延迟的表达式。具体来说,我们试图回答四个问题:1)复制和编码冗余在成本与延迟权衡方面如何比较?2)我们能否在等待一段时间后引入冗余,并期望它能降低成本?3)重新启动在一段时间后似乎落后的任务是否有助于降低成本和/或延迟?4)同时使用冗余和重新启动是否有效?我们通过使用从Google聚类数据中提取的经验分布进行模拟,验证了我们为每个问题找到的答案。
Runtime performance variability has been a major issue, hindering predictable and scalable performance in modern distributed systems. Executing requests or jobs redundantly over multiple servers have been shown to be effective for mitigating variability, both in theory and practice. Systems that employ redundancy has drawn significant attention, and numerous papers have analyzed the pain and gain of redundancy under various service models and assumptions on the runtime variability. This paper presents a cost (pain) vs. latency (gain) analysis of executing jobs of many tasks by employing replicated or erasure coded redundancy. The tail heaviness of service time variability is decisive on the pain and gain of redundancy and we quantify its effect by deriving expressions for cost and latency. Specifically, we try to answer four questions: 1) How do replicated and coded redundancy compare in the cost vs. latency tradeoff? 2) Can we introduce redundancy after waiting some time and expect it to reduce the cost? 3) Can relaunching the tasks that appear to be straggling after some time help to reduce cost and/or latency? 4) Is it effective to use redundancy and relaunching together? We validate the answers we found for each of these questions via simulations that use empirical distributions extracted from a Google cluster data.