Classification Based Metadata Management for HDFS

Classification Based Metadata Management for HDFS
复制标题

DOI:
10.1109/hpcc.2012.149
复制
发表时间:
2012-06
期刊:
2012 IEEE 14th International Conference on High Performance Computing and Communication & 2012 IEEE 9th International Conference on Embedded Software and Systems
影响因子:
--
通讯作者:
A. Chandrasekar;K. Chandrasekar;Harini Ramasatagopan;R. A. Rahim;J. Balasubramaniyan
A. Chandrasekar;K. Chandrasekar;Harini Ramasatagopan;R. A. Rahim;J. Balasubramaniyan
中科院分区:
其他
文献类型:
--
作者:
A. Chandrasekar;K. Chandrasekar;Harini Ramasatagopan;R. A. Rahim;J. Balasubramaniyan

文献摘要

被引文献

相似文献

查看数据存储的方式一直在变化。目前数据存储的趋势是Hadoop,它提供了一种可扩展的数据存储机制,用于存储海量数据和处理数据密集型的科学应用。它利用MapReduce框架,并将数据存储在HDFS(Hadoop分布式文件系统)中。在HDFS架构中,元数据由NameNode处理。在本文中,我们提出了一种新的有效的元数据动态管理机制,该机制基于元数据的重要因子(If)对元数据进行分类,If是数据的关键程度、访问频率和客户端使用数据的重要性的度量。元数据管理根据其重要性分为三种不同的技术。为了避免元数据成为NameNode主内存的约束,使用了序列文件的概念。这种方法带来了更高效的低延迟元数据操作,同时减少了NameNode主内存的瓶颈。
The way data storage is viewed has been changing consistently. The current trend in data storage is Hadoop which provides a scalable data storage mechanism for storing extremely large amount of data and to handle data intensive scientific applications. It makes use of the MapReduce framework and stores the data in HDFS(Hadoop Distributed File System). In HDFS architecture metadata is handled by the NameNode. In this paper, we propose a novel and efficient mechanism for managing the metadata dynamically by classifying the metadata based on its Importance Factor(If) which is a measure of the data's criticality, frequency of access and the importance of the client using the data. Metadata management is divided into three different techniques based on the importance. To save the amount of metadata from being a constraint on the main memory of the NameNode, the concept of sequence files is employed. This approach leads to more efficient low latency metadata operations, at the same time reduces the bottleneck of the NameNode main memory.