A Heterogeneity-Aware Task Scheduler for Spark

A Heterogeneity-Aware Task Scheduler for Spark
复制标题

DOI:
10.1109/cluster.2018.00042
复制
发表时间:
2018-09
期刊:
2018 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Luna Xu;A. Butt;Seung-Hwan Lim;R. Kannan
Luna Xu;A. Butt;Seung-Hwan Lim;R. Kannan
中科院分区:
其他
文献类型:
--
作者:
Luna Xu;A. Butt;Seung-Hwan Lim;R. Kannan

文献摘要

相似文献

像Spark这样的大数据处理系统被越来越多的不同应用所采用,例如机器学习、图形计算和科学计算,每个应用都有动态和不同的资源需求。这些应用程序越来越多地运行在异构硬件上,例如,用核外加速器然而,大数据平台并不考虑应用程序和硬件的多维异构性。这导致应用程序和硬件特性之间的根本不匹配,以及大数据平台中采用的资源调度。例如,Hadoop和Spark在将任务分配给节点时只考虑数据局部性,而通常忽略硬件功能和对特定应用程序需求的适用性。在本文中,我们提出了RUPAM,异构感知的大数据平台的任务调度系统,它考虑了任务级的资源特性和底层硬件特性,以及保持数据的局部性。RUPAM采用简单而有效的试探法来决定主导调度因素(例如,CPU、内存或I/O),在特定阶段中给定任务。我们的实验表明,RUPAM是能够提高代表性的应用程序的性能高达62.3%相比,标准的火花调度。
Big data processing systems such as Spark are employed in an increasing number of diverse applications—such as machine learning, graph computation, and scientific computing—each with dynamic and different resource needs. These applications increasingly run on heterogeneous hardware, e.g., with out-of-core accelerators. However, big data platforms do not factor in the multi-dimensional heterogeneity of applications and hardware. This leads to a fundamental mismatch between the application and hardware characteristics, and the resource scheduling adopted in big data platforms. For example, Hadoop and Spark consider only data locality when assigning tasks to nodes, and typically disregard the hardware capabilities and suitability to specific application requirements. In this paper, we present RUPAM, a heterogeneity-aware task scheduling system for big data platforms, which considers both task-level resource characteristics and underlying hardware characteristics, as well as preserves data locality. RUPAM adopts a simple yet effective heuristic to decide the dominant scheduling factor (e.g., CPU, memory, or I/O), given a task in a particular stage. Our experiments show that RUPAM is able to improve the performance of representative applications by up to 62.3% compared to the standard Spark scheduler.