Key based data analytics across data centers considering bi-level resource provision in cloud computing

Key based data analytics across data centers considering bi-level resource provision in cloud computing
复制标题

考虑云计算中双层资源提供的跨数据中心的基于密钥的数据分析

DOI:
10.1016/j.future.2016.03.008
复制
发表时间:
2016-09
期刊:
Future Generation Comp. Syst.
影响因子:
--
通讯作者:
Xuan Wang
Xuan Wang
中科院分区:
其他
文献类型:
--
作者:
Jiangtao Zhang;Lingmin Zhang;Hejiao Huang;Zeo L. Jiang;Xuan Wang

文献摘要

参考文献

被引文献

相似文献

由于数据源的分布特征,例如天文学和销售,或者法律禁止,将全球数据仅存储在一个数据中心(DC)中并不总是可行的。 Hadoop 是普遍接受的大数据分析框架。但它只能处理一个DC内的数据。数据的分布需要研究跨DC的Hadoop。不过,在这种情况下,我们可以将mappers放在本地DC中,在哪里放置reducer是一个很大的挑战,因为每个reducer需要处理所有涉及的DC上的几乎所有map输出。本文提出了一种新颖的架构和基于密钥的方案,该方案可以尽可能尊重传统Hadoop的局部性原则,同时实现以较低成本部署reducer。考虑到DC级和服务器级资源提供,采用双层编程对问题进行形式化,并通过定制的双层组遗传算法(TLGGA)进行求解。最终的结果可能分散在多个DC中,但可以聚合到指定的DC或传输和存储成本最小的DC。大量的模拟证明了 TLGGA 的有效性。它的性能分别比基线和最先进的机制高出 49% 和 40%。
Due to the distribution characteristic of the data source, such as astronomy and sales, or the legal prohibition, it is not always practical to store the world-wide data in only one data center (DC). Hadoop is a commonly accepted framework for big data analytics. But it can only deal with data within one DC. The distribution of data necessitates the study of Hadoop across DCs. In this situation, though, we can place mappers in the local DCs, where to place reducers is a great challenge, since each reducer needs to process almost allmapoutput across all involved DCs. In this paper, a novel architecture and akeybased scheme are proposed which can respect the locality principle of traditional Hadoop as much as possible while realizing deployment of reducers with lower costs. Considering both the DC level and the server level resource provision, bi-level programming is used to formalize the problem and it is solved by a tailored two level group genetic algorithm (TLGGA). The final results, which may be dispersed in several DCs, can be aggregated to a designative DC or the DC with the minimum transfer and storage cost. Extensive simulations demonstrate the effectiveness of TLGGA. It can outperform both the baseline and the state-of-the-art mechanisms by 49% and 40%, respectively.
DOI: --
发表时间: 2009-05
期刊: --
影响因子: --
作者:
Tom White
通讯作者: Tom White
DOI: 10.1007/978-3-642-54927-4_97
发表时间: 2014
期刊: --
影响因子: --
作者:
Peng Xu;H. Wang;Ming Tian
通讯作者: Peng Xu;H. Wang;Ming Tian
DOI: 10.1007/s10586-014-0359-y
发表时间: 2015-03
期刊: Cluster Computing
影响因子: --
作者:
F. F. Moghaddam-F.;R. F. Moghaddam;M. Cheriet
通讯作者: F. F. Moghaddam-F.;R. F. Moghaddam;M. Cheriet
DOI: 10.1002/spe.2229
发表时间: 2014-07
期刊: Software: Practice and Experience
影响因子: --
作者:
Wanfeng Zhang;Lizhe Wang;Yan Ma;Dingsheng Liu
通讯作者: Wanfeng Zhang;Lizhe Wang;Yan Ma;Dingsheng Liu
DOI: 10.1145/2287016.2287019
发表时间: 2012-06
期刊: --
影响因子: --
作者:
R. Tudoran;Alexandru Costan;Gabriel Antoniu
通讯作者: R. Tudoran;Alexandru Costan;Gabriel Antoniu