An Innovative Strategy for Improved Processing of Small Files in Hadoop

An Innovative Strategy for Improved Processing of Small Files in Hadoop
复制标题

DOI:
--
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
Priyanka G. Phakade;Dr. Suhas Raut
Priyanka G. Phakade;Dr. Suhas Raut
中科院分区:
其他
文献类型:
--
作者:
Priyanka G. Phakade;Dr. Suhas Raut

文献摘要

被引文献

相似文献

在互联网使用日益增多的今天,用户希望将数据存储在云计算平台上。大多数情况下,用户的数据都是小文件。HDFS旨在处理大量数据。但它不能处理大量的小文件。本文设计了一种改进的小文件处理模型。在现有系统中,映射任务一次处理一块输入。映射任务产生中间输出,该中间输出被提供给减速器。Reducer提供排序合并输出。这里使用多个映射任务和单个归约任务来处理小文件。在这种方法中,如果有大量的小文件,则每个映射任务获得的输入较少。这样,HDFS处理大量小文件的性能就下降了。在所提出的系统中,小文件被HDFS高效地处理。HDFS客户端请求在HDFS中存储小文件。NameNode允许在HDFS中存储小文件。当HDFS客户端提交小文件进行处理时,NameNode将文件合并为单个拆分,该拆分将成为映射任务的输入。将MAP任务产生的中间输出作为输入提供给多个减速器。Reducer提供排序的合并输出。随着MAP任务数量的减少,处理时间也会减少。
Nowadays, the use of internet grows, so user wish to store data on cloud computing platform. Most of the time, user’s data are small files. HDFS designed to process large volume of data. But it cannot handle large number of small files. In this paper, we have designed improved model for processing small files. In existing system, map tasks processes a block of input at a time. Map task produces intermediate output which is given to reducer. Reducer gives sort-merge output. Here, multiple map tasks & single reduce task is used for processing small file. In this approach, if there are large numbers of small files then each map task gets less input. In this way, performance of HDFS for processing lot of small files has been degraded. In proposed system, small files efficiently processed by HDFS. HDFS client requested to store small files in HDFS. NameNode permits to store small files in HDFS. When small files are submitted by HDFS client for processing, NameNode combine files into single split which becomes an input to map task. Intermediate output produced by map task is given to multiple reducers as an input. Reducer gives sorted merge output. As the number of map tasks has reduced, the processing time gets decreases.