Optimizing small file storage process of the HDFS which based on the indexing mechanism

Optimizing small file storage process of the HDFS which based on the indexing mechanism
复制标题

DOI:
10.1109/icccbda.2017.7951882
复制
发表时间:
2017-04
期刊:
2017 IEEE 2nd International Conference on Cloud Computing and Big Data Analysis (ICCCBDA)
影响因子:
--
通讯作者:
Wenjuan Cheng;Miaomiao Zhou;Bing Tong;Junhong Zhu
Wenjuan Cheng;Miaomiao Zhou;Bing Tong;Junhong Zhu
中科院分区:
其他
文献类型:
--
作者:
Wenjuan Cheng;Miaomiao Zhou;Bing Tong;Junhong Zhu

文献摘要

被引文献

相似文献

Hadoop分布式文件系统(HDFS)作为GFS的一种开源实现,在处理大文件时具有很高的效率。但由于其自身的主从式结构和元数据的存储方式,在处理海量小文件时效率较低。它占用NameNode大量内存,降低访问效率,延迟并发用户访问。为了提高这一性能效率,本文研究了在HDFS上处理小文件的方法。根据文件存储过程,提出了一种基于索引机制的小文件处理方案。在将文件上传到HDFS集群之前,会测量文件大小。小文件被编入索引并合并。如果它是一个小文件,那么它将被索引和处理。并创建一个索引文件来保存小文件的索引信息。同时,该方案引入分布式缓存策略,进一步优化小文件的I/O操作,从而提高阅读速度。实验结果表明,与原HDFS和HAR方案相比,该方案在获取内存效率和内存资源消耗方面有很大的改善。
As an open source implementation of GFS, Hadoop Distributed File System (HDFS) has high efficiency on handling the large files. However, due to its own master-slave structure and the storage of metadata, the efficiency is low when dealing with massive small files. It occupies large amount of NameNode memory, reduces access efficiency, and delays concurrent user access. In order to improve this performance efficiency, this paper studies the method of processing small files on HDFS. According to the file storage process, this paper proposes a small file processing scheme based on index mechanism. Before the file is uploaded to the HDFS cluster, the file size is measured. The small files are indexed and merged. If it is a small file, then it will be indexed and processed. And it will be created an index file to save the index information of the small file. At the same time, this scheme introduces the distributed caching strategy to further optimize the I/O operation of small files, so as to improve the reading speed. Experimental results show that compared with the original HDFS and HAR scheme, this scheme has a great improvement in obtain memory efficiency and consumption of memory resources.