Improving the Efficiency of Storing for Small Files in HDFS

Improving the Efficiency of Storing for Small Files in HDFS
复制标题

DOI:
10.1109/csss.2012.556
复制
发表时间:
2012-08
期刊:
2012 International Conference on Computer Science and Service System
影响因子:
--
通讯作者:
Yang Zhang;Dan Liu
Yang Zhang;Dan Liu
中科院分区:
其他
文献类型:
--
作者:
Yang Zhang;Dan Liu

文献摘要

被引文献

相似文献

HDFS(Hadoop分布式文件系统)是一种流行的文件系统。但是HDFS对于小文件有效率低下的问题。传统方法存在资源消耗大、效率低的缺点。为了解决这个问题,本文提出了一种新的小文件处理方法,它作为一个引擎独立于HDFS。该引擎可以有效降低HDFS的开销。它使用Reactor多路复用IO来构建服务器,并使用非阻塞IO来合并和读取小文件。并且该引擎有一个小文件的缓存,可以使阅读效率高。提出了一种小文件处理策略,通过建立文件索引,利用边界文件块填充机制实现文件分离和文件检索,实现文件的高效合并。实验结果表明,该方法提高了HDFS中存储和处理大量小文件的效率。
HDFS (Hadoop Distributed File System) is the popular file system. But HDFS has inefficient issue with small files. Traditional method has the drawback of high resource consumption and low efficiency performance. In order to resolve this problem, this paper proposes a novel approach for small files process, which works as an engine independent with the HDFS. This engine can reduce the overhead of HDFS effectively. It uses Reactor multiplexed IO to build the server and uses non-blocking IO to merge and read small files. And the engine has a cache of small files that can make the reading efficiently. This paper presents the small files processing strategy for files efficient merger, which builds the file index and uses boundary file block filling mechanism to accomplish files separation and files retrieval. At last the experimental results show that the novel approach has improved the efficiency of storing and processing massive small files in HDFS.