FlashR: parallelize and scale R for machine learning using SSDs

FlashR: parallelize and scale R for machine learning using SSDs
复制标题

FlashR:使用 SSD 并行化和扩展 R 以进行机器学习

DOI:
10.1145/3200691.3178501
复制
发表时间:
2018
影响因子:
--
通讯作者:
Burns, Randal
Burns, Randal
中科院分区:
--
文献类型:
--
作者:
Zheng, Da;Mhembere, Disa;Vogelstein, Joshua T.;Priebe, Carey E.;Burns, Randal

文献摘要

参考文献

相似文献

R是用于统计和机器学习的最流行的编程语言之一,但它速度很慢,无法扩展到大型数据集。在R中实现高效算法的一般方法是用C或FORTRAN实现它,并提供一个R包装器。FlashR通过并行化Rbasepackage中的大量矩阵函数,并利用固态硬盘(SSD)将它们扩展到超出内存容量,从而加速和扩展现有的R代码。FlashR通过(I)惰性地评估矩阵操作,(Ii)在一次执行中执行DAG中的所有操作并且只有一次遍历数据以增加计算与I/O的比率,(Iii)在矩阵分区上执行两级矩阵划分和重新排序计算以减少存储器层次中的数据移动,来执行存储器层次感知执行以加速并行化的R代码。我们在各种机器学习和统计算法上对FlashR进行了评估,这些算法的输入多达40亿个数据点。尽管固态硬盘和内存之间存在巨大的性能差距,但固态硬盘上的FlashR对于许多算法都密切跟踪内存中的FlashR的性能。FlashR中的R实现比H2O和Spark MLlib的性能高出3-20倍。
R is one of the most popular programming languages for statistics and machine learning, but it is slow and unable to scale to large datasets. The general approach for having an efficient algorithm in R is to implement it in C or FORTRAN and provide an R wrapper. FlashR accelerates and scales existing R code by parallelizing a large number of matrix functions in the Rbasepackage and scaling them beyond memory capacity with solid-state drives (SSDs). FlashR performs memory hierarchy aware execution to speed up parallelized R code by(i)evaluating matrix operations lazily,(ii)performing all operations in a DAG in a single execution and with only one pass over data to increase the ratio of computation to I/O,(iii)performing two levels of matrix partitioning and reordering computation on matrix partitions to reduce data movement in the memory hierarchy. We evaluate FlashR on various machine learning and statistics algorithms on inputs of up to four billion data points. Despite the huge performance gap between SSDs and RAM, FlashR on SSDs closely tracks the performance of FlashR in memory for many algorithms. The R implementations in FlashR outperforms H2O and Spark MLlib by a factor of 3 -- 20.
DOI: --
发表时间: 2012
影响因子: 1.5
作者:
Wai;Da Zheng
通讯作者: Da Zheng
DOI: --
发表时间: 2022
期刊:
影响因子: --
作者:
Tsuchikawa T.;Kaneda H.;Oyabu S.;Kokusho T.;Kobayashi H.;Yamagishi M.;Toba Y.;Kiyoto Yoshino
通讯作者: Kiyoto Yoshino
Riposte:R 中矢量代码的跟踪驱动编译器和并行 VM
DOI: --
发表时间: 2012
期刊: International Conference on Parallel Architectures and Compilation Techniques
影响因子: --
作者:
Justin Talbot;Zach DeVito;P. Hanrahan
通讯作者: P. Hanrahan
APL 中的编译和延迟评估
DOI: --
发表时间: 1978
期刊: ACM-SIGACT Symposium on Principles of Programming Languages
影响因子: --
作者:
L. Guibas;Douglas K. Wyatt
通讯作者: Douglas K. Wyatt
DESOLA:使用延迟评估和运行时代码生成的主动线性代数库
DOI: --
发表时间: 2011
影响因子: 1.3
作者:
F. P. Russell;Michael R. Mellor;P. Kelly;Olav Beckmann
通讯作者: Olav Beckmann