Parallelizing R in Hadoop (A Work-in-Progress Study)

Parallelizing R in Hadoop (A Work-in-Progress Study)
复制标题

在 Hadoop 中并行化 R(一项正在进行的研究)

DOI:
--
复制
发表时间:
2015
期刊:
International Conference on Smart Cities
影响因子:
--
通讯作者:
Hung
Hung
中科院分区:
--
文献类型:
--
作者:
Yen;Yu;Chia;Hung

文献摘要

被引文献

相似文献

R 是一种流行的编程语言,被数据科学家广泛采用。然而,典型的R只能在单机环境中执行。虽然R可以链接到Hadoop(例如RHadoop),但R用户需要基于MapReduce框架来开发R脚本。这要求 R 程序员具有很高的技能,以在 Map 和 Reduce 作业方面并行化他们的 R 程序,从而扼杀了在分布式环境中执行 R 计算超过单机容量的动力。在本文中,我们提出了一种在 Hadoop 中并行化 R 的实现。我们的目标是让 R 用户无需修改即可在 Hadoop 中运行在单机环境中开发的 R 脚本。虽然这项研究工作仍在进行中,但我们在本文中报告了我们关于如何隐藏在 Hadoop 中迁移和运行此类 R 脚本的复杂性的初步经验。
R is a popular programming language which is widely adopted by data scientists. However, typical R can only be executed in a single machine environment. Although R can be linked to Hadoop such as RHadoop, R users need to develop their R scripts based on the MapReduce framework. This demands highly skill of R programmers to parallelize their R pro-grams in terms of Map and Reduce jobs, killing the motivation of performing R computation in distributed environments out-pacing the single machine capacity. We present an implementation for parallelizing R in Hadoop in this paper. Our objective is to allow R users to run their R scripts, which are developed in a single machine environment, in Hadoop without modification. While this research work is still ongoing, we report our preliminary experiences in this paper on how to hide the complexity of migrating and running such R scripts in Hadoop.