Parallelizing R in Hadoop (A Work-in-Progress Study)
Parallelizing R in Hadoop (A Work-in-Progress Study)
复制标题
在 Hadoop 中并行化 R(一项正在进行的研究)
DOI:
--
复制
发表时间:
2015
期刊:
影响因子:
--
通讯作者:
Hung
中科院分区:
文献类型:
--
作者:
Yen;Yu;Chia;Hung
R is a popular programming language which is widely adopted by data scientists. However, typical R can only be executed in a single machine environment. Although R can be linked to Hadoop such as RHadoop, R users need to develop their R scripts based on the MapReduce framework. This demands highly skill of R programmers to parallelize their R pro-grams in terms of Map and Reduce jobs, killing the motivation of performing R computation in distributed environments out-pacing the single machine capacity. We present an implementation for parallelizing R in Hadoop in this paper. Our objective is to allow R users to run their R scripts, which are developed in a single machine environment, in Hadoop without modification. While this research work is still ongoing, we report our preliminary experiences in this paper on how to hide the complexity of migrating and running such R scripts in Hadoop.