ARRA: An Associated Replica Replacement Algorithm Based on Apriori Approach for Data Intensive Jobs in Data Grid

ARRA: An Associated Replica Replacement Algorithm Based on Apriori Approach for Data Intensive Jobs in Data Grid
复制标题

DOI:
10.4028/www.scientific.net/kem.439-440.1409
复制
发表时间:
2010-06
期刊:
Key Engineering Materials
影响因子:
--
通讯作者:
J. Jiang;H. Ji;Gaochao Xu;Xiaohui Wei
J. Jiang;H. Ji;Gaochao Xu;Xiaohui Wei
中科院分区:
其他
文献类型:
--
作者:
J. Jiang;H. Ji;Gaochao Xu;Xiaohui Wei

文献摘要

相似文献

在数据网格中,创建多个副本是处理数据密集型作业的一种有效策略。更换电池是这一战略的关键一步。为了解决副本替换问题,人们提出了经济模型、流行度模型和混合模型等基于每个数据文件进行分析和预测,但这些模型忽略了不同数据文件之间的关联关系。为了发现隐藏在数据密集型作业中的这些关联关系,采用数据挖掘领域的Apriori算法分析每个数据密集型作业的行为。提出了一种基于Apriori方法的数据网格关联副本替换算法。该算法分为两个主要步骤:1)对各节点上的数据文件进行关联行为分析和分类; 2)副本替换规则的生成和应用。在Optorsim中对该算法进行了仿真,并与LFU算法进行了比较。实验结果表明,与LFU相比,在所有作业的平均作业时间、远程文件访问次数和有效的网络使用率方面具有相对优势。
Creating many replicas in the processing of data-intensive jobs in data grid is an efficient strategy. Replica replacement is the crucial step to this strategy. Economic model, popularity model and hybrid model etc. have been proposed to solve this issue of replica replacement with analysis and prediction based on each data file, however, these models neglect association relationships among different data files. To find out these association relationships hidden in data-intensive jobs, Apriori algorithm in data mining field is adopted to analyze behaviors of each data-intensive job. An associated replica replacement algorithm based on Apriori approach in data grid is proposed in this paper. This algorithm has two major steps: 1) associated behavior analysis and classification of data files in each node; 2) generation and application of replica replacement rules. Our proposed algorithm is simulated in Optorsim to be compared with LFU algorithm. The experiment shows that there is a relative advantage compared with LFU in mean job times of all jobs, number of remote file access and effective network usage perspectives.