A novel indexing scheme for efficient handling of small files in Hadoop Distributed File System

A novel indexing scheme for efficient handling of small files in Hadoop Distributed File System
复制标题

DOI:
10.1109/iccci.2013.6466147
复制
发表时间:
2013-02
期刊:
2013 International Conference on Computer Communication and Informatics
影响因子:
--
通讯作者:
S. Chandrasekar;R. Dakshinamurthy;P G Seshakumar;B. Prabavathy;C. Babu
S. Chandrasekar;R. Dakshinamurthy;P G Seshakumar;B. Prabavathy;C. Babu
中科院分区:
其他
文献类型:
--
作者:
S. Chandrasekar;R. Dakshinamurthy;P G Seshakumar;B. Prabavathy;C. Babu

文献摘要

被引文献

相似文献

Hadoop分布式文件系统(HDFS)旨在可靠地存储和管理超大文件。HDFS中的所有文件都由单个服务器NameNode管理。NameNode在其主内存中存储存储到HDFS中的每个文件的元数据。因此,HDFS会因小文件数量的增加而遭受性能损失。存储和管理大量的小文件给NameNode带来了沉重的负担。可以存储到HDFS中的文件数量受到NameNode主内存大小的限制。此外,HDFS没有考虑文件之间的相关性,也没有提供任何预取机制来提高I/O性能。为了提高HDFS上小文件的存储和访问效率,在Dong等人工作的基础上提出了一种解决方案,扩展Hadoop分布式文件系统(EHDFS)。在这种方法中,一组相关的文件被组合成单个大文件,以减少文件数量。建立了一个索引机制,从相应的组合文件中访问各个文件。此外,还提供了索引预取以提高I/O性能并最小化NameNode上的负载。实验结果表明,EHDFS能够减少NameNode主内存上的元数据占用量16%,并提高存储和访问大量小文件的效率。
Hadoop Distributed File System (HDFS) is designed for reliable storage and management of very large files. All the files in HDFS are managed by a single server, the NameNode. NameNode stores metadata, in its main memory, for each file stored into HDFS. As a consequence, HDFS suffers a performance penalty with increased number of small files. Storing and managing a large number of small files imposes a heavy burden on the NameNode. The number of files that can be stored into HDFS is constrained by the size of NameNode's main memory. Further, HDFS does not take the correlation among files into account, and it does not provide any prefetching mechanism to improve the I/O performance. In order to improve the efficiency of storing and accessing the small files on HDFS, we propose a solution based on the works of Dong et al., namely Extended Hadoop Distributed File System (EHDFS). In this approach, a set of correlated files is combined, as identified by the client, into a single large file to reduce the file count. An indexing mechanism has been built to access the individual files from the corresponding combined file. Further, index prefetching is also provided to improve I/O performance and minimize the load on NameNode. The experimental results indicate that EHDFS is able to reduce the metadata footprint on NameNode's main memory by 16% and also improve the efficiency of storing and accessing large number of small files.