Parallel data intensive applications using MapReduce: a data mining case study in biomedical sciences

Parallel data intensive applications using MapReduce: a data mining case study in biomedical sciences
复制标题

DOI:
10.1007/s10586-014-0405-9
复制
发表时间:
2015-03
期刊:
Cluster Computing
影响因子:
--
通讯作者:
Liangxiu Han;Hwee Yong Ong
Liangxiu Han;Hwee Yong Ong
中科院分区:
其他
文献类型:
--
作者:
Liangxiu Han;Hwee Yong Ong

文献摘要

被引文献

相似文献

性能在数据密集型应用程序(例如数据挖掘任务)中是一个开放的问题。并行和分布式计算系统(如多核计算、网格计算、云计算等),以及混合编程模型(如MapReduce、MPI等),被视为加速数据密集型应用程序的热门解决方案。主要挑战之一是如何有效地利用这些先进技术来促进诸如生物医学科学等基础科学发现。本文探讨了MapReduce和云计算如何通过生物医学科学中的真实数据挖掘用例来加速数据密集型应用程序的性能。我们首先使用MapReduce模型调整数据挖掘任务,然后将其部署到云上。基于MapReduce的计算,我们建立了一个分析模型来评估原型的效率和性能。实验和评估模型的结果表明,通过这些先进技术可以提高性能和可扩展性。
Performance is an open issue in data intensive applications (e.g. data mining tasks). Parallel and distributed computing systems (e.g. multicore computing, grid computing, cloud computing,etc.), along with hybrid programming models (e.g. MapReduce, MPI, etc.), is seen a sought-after solution for accelerating data-intensive applications. One of main challenges is how to exploit these advanced technologies effectively in facilitating fundamental science discoveries such as those in Biomedical Sciences. This paper explores how MapReduce and Cloud computing can accelerate performance of data intensive applications through a real data mining use case in the Biomedical Sciences. We have first adapted the data mining task using MapReduce model and then deployed it onto the Cloud. We have built an analytic model based on the MapReduce computations to evaluate the efficiency and performance of the prototype. The results, from both experiments and the evaluation model, show the performance and scalability can be enhanced through these advanced technologies.