Load Rebalancing for Distributed File Systems in Clouds

Load Rebalancing for Distributed File Systems in Clouds
复制标题

DOI:
10.1109/tpds.2012.196
复制
发表时间:
2013-05-01
影响因子:
5.3
通讯作者:
Chao, Yu-Chang
Chao, Yu-Chang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hsiao, Hung-Chang;Chung, Hsueh-Yi;Chao, Yu-Chang

文献摘要

被引文献

相似文献

分布式文件系统是基于MapReduce编程范式的云计算应用程序的关键构建块。在这种文件系统中,节点同时具有计算和存储功能;一个文件被划分为多个块,这些块被分配到不同的节点上,这样MapReduce任务就可以在节点上并行执行。但是,在云计算环境中,故障是常态,并且可能会升级、替换和添加节点。还可以动态地创建、删除和追加文件。这将导致分布式文件系统的负载不平衡;也就是说,文件块在节点之间的分布并不均匀。生产系统中出现的分布式文件系统强烈依赖于中心节点进行块重新分配。在大规模、容易发生故障的环境中,这种依赖关系显然是不够的,因为中央负载平衡器的工作负载与系统大小呈线性关系,因此可能成为性能瓶颈和单点故障。本文提出了一种完全分布式的负载再平衡算法来解决负载不平衡问题。我们的算法与生产系统中的集中式方法和文献中提出的竞争性分布式解决方案进行了比较。仿真结果表明,我们的方法与现有的集中式方法相当,并且在负载不平衡因素、移动成本和算法开销方面明显优于先前的分布式算法。在集群环境下,进一步研究了在Hadoop分布式文件系统中实现的方案的性能。
Distributed file systems are key building blocks for cloud computing applications based on the MapReduce programming paradigm. In such file systems, nodes simultaneously serve computing and storage functions; a file is partitioned into a number of chunks allocated in distinct nodes so that MapReduce tasks can be performed in parallel over the nodes. However, in a cloud computing environment, failure is the norm, and nodes may be upgraded, replaced, and added in the system. Files can also be dynamically created, deleted, and appended. This results in load imbalance in a distributed file system; that is, the file chunks are not distributed as uniformly as possible among the nodes. Emerging distributed file systems in production systems strongly depend on a central node for chunk reallocation. This dependence is clearly inadequate in a large-scale, failure-prone environment because the central load balancer is put under considerable workload that is linearly scaled with the system size, and may thus become the performance bottleneck and the single point of failure. In this paper, a fully distributed load rebalancing algorithm is presented to cope with the load imbalance problem. Our algorithm is compared against a centralized approach in a production system and a competing distributed solution presented in the literature. The simulation results indicate that our proposal is comparable with the existing centralized approach and considerably outperforms the prior distributed algorithm in terms of load imbalance factor, movement cost, and algorithmic overhead. The performance of our proposal implemented in the Hadoop distributed file system is further investigated in a cluster environment.