Sparrow: distributed, low latency scheduling

Sparrow: distributed, low latency scheduling
复制标题

DOI:
10.1145/2517349.2522716
复制
发表时间:
2013-11
期刊:
Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles
影响因子:
--
通讯作者:
Kay Ousterhout;Patrick Wendell;M. Zaharia;I. Stoica
Kay Ousterhout;Patrick Wendell;M. Zaharia;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
Kay Ousterhout;Patrick Wendell;M. Zaharia;I. Stoica

文献摘要

被引文献

相似文献

大规模数据分析框架正在转向更短的任务持续时间和更大程度的并行性,以提供低延迟。调度在数百毫秒内完成的高度并行作业对任务调度程序提出了重大挑战,任务调度程序需要在适当的机器上每秒调度数百万个任务,同时提供毫秒级延迟和高可用性。我们证明,分散式随机采样方法可提供接近最佳的性能,同时避免集中式设计的吞吐量和可用性限制。我们在 110 台机器的集群上实现并部署了我们的调度程序 Sparrow,并证明 Sparrow 的性能与理想调度程序相差 12% 以内。
Large-scale data analytics frameworks are shifting towards shorter task durations and larger degrees of parallelism to provide low latency. Scheduling highly parallel jobs that complete in hundreds of milliseconds poses a major challenge for task schedulers, which will need to schedule millions of tasks per second on appropriate machines while offering millisecond-level latency and high availability. We demonstrate that a decentralized, randomized sampling approach provides near-optimal performance while avoiding the throughput and availability limitations of a centralized design. We implement and deploy our scheduler, Sparrow, on a 110-machine cluster and demonstrate that Sparrow performs within 12% of an ideal scheduler.