Classification Based Metadata Management for HDFS
Classification Based Metadata Management for HDFS
复制标题
DOI:
10.1109/hpcc.2012.149
复制
发表时间:
2012-06
期刊:
影响因子:
--
通讯作者:
A. Chandrasekar;K. Chandrasekar;Harini Ramasatagopan;R. A. Rahim;J. Balasubramaniyan
中科院分区:
文献类型:
--
作者:
A. Chandrasekar;K. Chandrasekar;Harini Ramasatagopan;R. A. Rahim;J. Balasubramaniyan
The way data storage is viewed has been changing consistently. The current trend in data storage is Hadoop which provides a scalable data storage mechanism for storing extremely large amount of data and to handle data intensive scientific applications. It makes use of the MapReduce framework and stores the data in HDFS(Hadoop Distributed File System). In HDFS architecture metadata is handled by the NameNode. In this paper, we propose a novel and efficient mechanism for managing the metadata dynamically by classifying the metadata based on its Importance Factor(If) which is a measure of the data's criticality, frequency of access and the importance of the client using the data. Metadata management is divided into three different techniques based on the importance. To save the amount of metadata from being a constraint on the main memory of the NameNode, the concept of sequence files is employed. This approach leads to more efficient low latency metadata operations, at the same time reduces the bottleneck of the NameNode main memory.