Network-aware resource management for scalable data analytics frameworks

Network-aware resource management for scalable data analytics frameworks
复制标题

DOI:
10.1109/bigdata.2015.7364083
复制
发表时间:
2015-10
期刊:
2015 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
T. Renner;L. Thamsen;O. Kao
T. Renner;L. Thamsen;O. Kao
中科院分区:
其他
文献类型:
--
作者:
T. Renner;L. Thamsen;O. Kao

文献摘要

相似文献

在多个框架、应用程序和数据集之间共享集群资源对于进行大规模数据分析的组织非常重要。它提高了集群利用率,避免了仅运行单个框架的独立集群,并允许数据科学家为每个分析任务选择最佳框架。当前用于集群资源管理的系统(如YARN或Mesos)使用容器来实现资源共享。分析框架在这些容器中执行任务。然而,目前的容器放置主要基于内核和存储器方面的可用计算能力,但忽略了也考虑网络拓扑和数据位置。在本文中,我们提出了一种容器放置方法,(a)考虑到网络拓扑结构,以防止网络的核心网络和(B)的地方容器靠近输入数据,以提高数据的本地性,减少远程磁盘读取分布式文件系统。在容器放置级别上引入拓扑和数据感知的主要优点是,多个应用程序框架可以从改进中受益。我们提出了一个与Hadoop YARN集成的原型,并使用Apache Flink对由不同应用程序和数据集组成的工作负载进行了评估。我们的评估64核心集群,其中节点连接通过胖树拓扑结构,显示出可喜的结果与加速高达67%的网络密集型工作负载。
Sharing cluster resources between multiple frameworks, applications and datasets is important for organizations doing large scale data analytics. It improves cluster utilization, avoids standalone clusters running only a single framework and allows data scientists to choose the best framework for each analysis task. Current systems for cluster resource management like YARN or Mesos achieve resource sharing using containers. Analytics frameworks execute their tasks in these containers. However, currently the container placement is based predominantly on available computing capabilities in terms of cores and memory, yet neglects to also take the network topology and data locations into account. In this paper, we propose a container placement approach that (a) takes the network topology into account to prevent network congestions in the core network and (b) places containers close to input data to improve data locality and reduce remote disk reads in distributed file systems. The main advantages of introducing topology- and data-awareness on the level of container placement is that multiple application frameworks benefit from improvements. We present a prototype integrated with Hadoop YARN and an evaluation with workloads consisting of different applications and datasets using Apache Flink. Our evaluation on a 64 core cluster, in which nodes are connected through a fat tree topology, shows promising results with speedups of up to 67% for network-intensive workloads.