A Study of Effective Replica Reconstruction Schemes at Node Deletion for HDFS

A Study of Effective Replica Reconstruction Schemes at Node Deletion for HDFS
复制标题

DOI:
10.1109/ccgrid.2014.31
复制
发表时间:
2014-05
期刊:
2014 14th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing
影响因子:
--
通讯作者:
Asami Higai;A. Takefusa;H. Nakada;M. Oguchi
Asami Higai;A. Takefusa;H. Nakada;M. Oguchi
中科院分区:
其他
文献类型:
--
作者:
Asami Higai;A. Takefusa;H. Nakada;M. Oguchi

文献摘要

被引文献

相似文献

分布式文件系统在多个商用机器上管理大量数据,作为大数据应用的管理和处理系统已经引起了人们的关注。分布式文件系统由多个数据节点组成,并通过保存多个数据副本来提供可靠性和可用性。由于系统故障或维护,数据节点可能从系统中移除,并且移除的数据节点所保持的数据块丢失。如果数据块丢失,则保持丢失的数据块的其他数据节点的访问负载增加,并且结果,分布式文件系统上的数据处理的性能降低。因此,副本重建是一个重要的问题,重新分配丢失的数据块,以防止这种性能下降。Hadoop分布式文件系统(HDFS)是一种广泛使用的分布式文件系统。在HDFS副本重建过程中,随机选择用于复制的源和目的数据节点。我们发现,这种副本重建方案是低效的,因为数据传输是有偏见的。因此,我们提出了两个更有效的副本重建计划,旨在平衡复制过程的工作负载。我们提出的复制调度策略假设节点被安排在一个环和数据块传输基于这个单向环结构,以最大限度地减少每个节点的传输数据量的差异。基于此策略,我们提出了两个副本重建方案,一个优化方案和一个启发式方案。我们已经在HDFS中实现了所提出的方案,并在实际的HDFS集群上对其进行了评估。从实验中,我们证实,副本重建吞吐量的建议计划显示了45%的改善相比,默认方案。我们还验证了启发式方案是有效的,因为它显示出的性能相媲美的优化方案,可以比优化方案更具可扩展性。
Distributed file systems, which manage large amounts of data over multiple commercially available machines, have attracted attention as a management and processing system for big data applications. A distributed file system consists of multiple data nodes and provides reliability and availability by holding multiple replicas of data. Due to system failure or maintenance, a data node may be removed from the system and the data blocks the removed data node held are lost. If data blocks are missing, the access load of the other data nodes that hold the lost data blocks increases, and as a result the performance of data processing over the distributed file system decreases. Therefore, replica reconstruction is an important issue to reallocate the missing data blocks in order to prevent such performance degradation. The Hadoop Distributed File System (HDFS) is a widely used distributed file system. In the HDFS replica reconstruction process, source and destination data nodes for replication are selected randomly. We found that this replica reconstruction scheme is inefficient because data transfer is biased. Therefore, we propose two more effective replica reconstruction schemes that aim to balance the workloads of replication processes. Our proposed replication scheduling strategy assumes that nodes are arranged in a ring and data blocks are transferred based on this one-directional ring structure to minimize the difference of the amount of transfer data of each node. Based on this strategy, we propose two replica reconstruction schemes, an optimization scheme and a heuristic scheme. We have implemented the proposed schemes in HDFS and evaluated them on an actual HDFS cluster. From the experiments, we confirm that the replica reconstruction throughput of the proposed schemes show a 45% improvement compared to that of the default scheme. We also verify that the heuristic scheme is effective because it shows performance comparable to the optimization scheme and can be more scalable than the optimization scheme.