Sol: Fast Distributed Computation Over Slow Networks

Sol: Fast Distributed Computation Over Slow Networks
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Fan Lai;Jie You;Xiangfeng Zhu;H. Madhyastha;Mosharaf Chowdhury
Fan Lai;Jie You;Xiangfeng Zhu;H. Madhyastha;Mosharaf Chowdhury
中科院分区:
其他
文献类型:
--
作者:
Fan Lai;Jie You;Xiangfeng Zhu;H. Madhyastha;Mosharaf Chowdhury

文献摘要

相似文献

大数据和人工智能的普及导致了分布式计算堆栈不同层的许多优化。尽管--或者也许是因为--它是这类软件堆栈中的小角色,但负责执行作业每一项任务的执行引擎的设计基本上没有改变。因此,目前可用的执行引擎主要是为低延迟和高带宽数据中心网络设计的。当其中一个或两个网络假设都不成立时,fi明显得不到充分利用。在本文中,我们采用fiRST原理的方法来开发一个能够适应不同网络条件的执行引擎。SOL,我们的联合执行引擎架构,fl在两个方面的现状。首先,为了减轻高延迟的影响,SOL主动分配任务,但这样做是明智的,以应对不确定性。其次,为了提高整体资源利用率,SOL在内部将通信与计算分离,而不是将资源同时提交给任务的两个方面。我们在EC2上的评测表明,在资源受限的网络中,相比于ApacheSpark,SOL将SQL和机器学习任务分别提高了16.4倍和4.2倍。
The popularity of big data and AI has led to many optimizations at different layers of distributed computation stacks. Despite – or perhaps, because of – its role as the narrow waist of such software stacks, the design of the execution engine, which is in charge of executing every single task of a job, has mostly remained unchanged. As a result, the execution engines available today are ones primarily designed for low latency and high bandwidth datacenter networks. When either or both of the network assumptions do not hold, CPUs are significantly underutilized. In this paper, we take a first-principles approach toward developing an execution engine that can adapt to diverse network conditions. Sol, our federated execution engine architecture , flips the status quo in two respects. First, to mitigate the impact of high latency, Sol proactively assigns tasks, but does so judiciously to be resilient to uncertainties. Second, to improve the overall resource utilization, Sol decouples communication from computation internally instead of commit-ting resources to both aspects of a task simultaneously. Our evaluations on EC2 show that, compared to Apache Spark in resource-constrained networks, Sol improves SQL and machine learning jobs by 16.4 × and 4.2 × on average.