Riposte: A trace-driven compiler and parallel VM for vector code in R

Riposte: A trace-driven compiler and parallel VM for vector code in R
复制标题

Riposte:R 中矢量代码的跟踪驱动编译器和并行 VM

DOI:
--
复制
发表时间:
2012
期刊:
International Conference on Parallel Architectures and Compilation Techniques
影响因子:
--
通讯作者:
P. Hanrahan
P. Hanrahan
中科院分区:
--
文献类型:
--
作者:
Justin Talbot;Zach DeVito;P. Hanrahan

文献摘要

被引文献

相似文献

用于数据分析的现代硬件和现代编程语言之间的利用率差距越来越大。由于功率和其他限制,最近的处理器设计寻求通过增加SIMD和多核并行性来提高性能。与此同时,用于数据分析的高级、动态类型的语言也变得流行起来。这些语言强调易用性和高生产率,但通常性能较低,对利用硬件并行性的支持有限。在本文中,我们描述了一种新的R语言运行时Reposte,它弥补了这一差距。Riposte使用跟踪,这是一种通常用于加速标量代码的技术,以动态发现并从任意R代码中提取向量操作序列。提取后,我们可以融合轨迹以消除不必要的内存流量,编译它们以使用硬件SIMD单元,并调度它们跨多个内核运行,从而使我们能够充分利用现代共享内存机器上的可用并行性。我们的测试表明,Reposte运行向量R代码的速度接近手工优化C语言的速度,比R的开源实现快5-50倍,并且对于某些任务,还可以线性扩展到32核。在12个不同的工作负载中,我们在没有明确的程序员并行化的情况下实现了150倍以上的总体平均加速。
There is a growing utilization gap between modern hardware and modern programming languages for data analysis. Due to power and other constraints, recent processor design has sought improved performance through increased SIMD and multi-core parallelism. At the same time, high-level, dynamically typed languages for data analysis have become popular. These languages emphasize ease of use and high productivity, but have, in general, low performance and limited support for exploiting hardware parallelism. In this paper, we describe Riposte, a new runtime for the R language, which bridges this gap. Riposte uses tracing, a technique commonly used to accelerate scalar code, to dynamically discover and extract sequences of vector operations from arbitrary R code. Once extracted, we can fuse traces to eliminate unnecessary memory traffic, compile them to use hardware SIMD units, and schedule them to run across multiple cores, allowing us to fully utilize the available parallelism on modern shared-memory machines. Our evaluation shows that Riposte can run vector R code near the speed of hand-optimized C, 5–50× faster than the open source implementation of R, and can also linearly scale to 32 cores for some tasks. Across 12 different workloads we achieve an overall average speed-up of over 150× without explicit programmer parallelization.