An optimized approach for storing and accessing small files on cloud storage

An optimized approach for storing and accessing small files on cloud storage
复制标题

DOI:
10.1016/j.jnca.2012.07.009
复制
发表时间:
2012-11
期刊:
J. Netw. Comput. Appl.
影响因子:
--
通讯作者:
B. Dong;Q. Zheng;Feng Tian;K. Chao;Rui Ma;R. Anane
B. Dong;Q. Zheng;Feng Tian;K. Chao;Rui Ma;R. Anane
中科院分区:
其他
文献类型:
--
作者:
B. Dong;Q. Zheng;Feng Tian;K. Chao;Rui Ma;R. Anane

文献摘要

被引文献

相似文献

Hadoop分布式文件系统(HDFS)被广泛采用来支持Internet服务。不幸的是,原生HDFS对于大量但小尺寸的文件表现不佳,这引起了人们的极大关注。首先分析并指出了HDFS小文件问题产生的原因:(1)大量的小文件对HDFS的NameNode造成了沉重的负担;(2)数据放置没有考虑小文件之间的相关性;(3)没有提供预取等优化机制来提高I/O性能。其次,在HDFS的上下文中,大文件和小文件之间的明确分界点是通过实验确定的,这有助于确定“多小才算小”。然后,根据文件关联特征,将文件分为结构关联文件、逻辑关联文件和独立文件三种类型。最后,在上述三个步骤的基础上,设计了一种优化的方法来提高小文件在HDFS上的存储和访问效率。文件合并和预取方案适用于结构相关的小文件,而文件分组和预取方案用于管理逻辑相关的小文件。实验结果表明,与原生HDFS和Hadoop文件归档工具相比,该方案有效地提高了小文件的存储和访问效率.
Hadoop distributed file system (HDFS) is widely adopted to support Internet services. Unfortunately, native HDFS does not perform well for large numbers but small size files, which has attracted significant attention. This paper firstly analyzes and points out the reasons of small file problem of HDFS: (1) large numbers of small files impose heavy burden on NameNode of HDFS; (2) correlations between small files are not considered for data placement; and (3) no optimization mechanism, such as prefetching, is provided to improve I/O performance. Secondly, in the context of HDFS, the clear cut-off point between large and small files is determined through experimentation, which helps determine ‘how small is small’. Thirdly, according to file correlation features, files are classified into three types: structurally-related files, logically-related files, and independent files. Finally, based on the above three steps, an optimized approach is designed to improve the storage and access efficiencies of small files on HDFS. File merging and prefetching scheme is applied for structurally-related small files, while file grouping and prefetching scheme is used for managing logically-related small files. Experimental results demonstrate that the proposed schemes effectively improve the storage and access efficiencies of small files, compared with native HDFS and a Hadoop file archiving facility.