Gerenuk: thin computation over big native data using speculative program transformation

Gerenuk: thin computation over big native data using speculative program transformation
复制标题

DOI:
10.1145/3341301.3359643
复制
发表时间:
2019-10
期刊:
Proceedings of the 27th ACM Symposium on Operating Systems Principles
影响因子:
--
通讯作者:
Christian Navasca;Cheng Cai;Khanh Nguyen;Brian Demsky;Shan Lu;Miryung Kim;G. Xu
Christian Navasca;Cheng Cai;Khanh Nguyen;Brian Demsky;Shan Lu;Miryung Kim;G. Xu
中科院分区:
其他
文献类型:
--
作者:
Christian Navasca;Cheng Cai;Khanh Nguyen;Brian Demsky;Shan Lu;Miryung Kim;G. Xu

文献摘要

相似文献

由于它们提供的快速开发周期,通常以面向对象的语言(例如Java和Scala)实现大数据系统。这些系统是在托管运行时(例如Java Virtual Machine(JVM))执行的,该系统要求每个数据项在对象进行处理之前以对象的处理。这种表示是多种严重效率低下的直接原因。我们开发了Gerenuk,这是一种编译器和运行时,旨在通过转换系统中的一组语句来实现基于JVM的数据并行系统,以实现近乎本地的效率,以直接执行与内线的本机BYTE。导致Gerenuk成功的关键见解是两个方面:(1)分析工作负载经常使用不可变和限制的数据类型。如果我们使用此假设对系统和用户代码进行了推测优化,则可以使转换可进行处理。 (2)数据流从一个值点开始,其中从本机字节和结束的序列化点创建对象,在序列化点,它们被转回到磁盘或网络的字节序列中。该流量自然定义了要转换的投机执行区(SER)。 Gerenuk将SER推定为可以直接通过来自磁盘或网络的本机字节操作的版本。 Gerenuk运行时在违反不变性和限制假设的情况下取消了SER的执行,并通过对字节进行反序列化并重新执行原始SER来切换到慢速路径。我们对Spark和Hadoop的评估证明了令人鼓舞的结果。
Big Data systems are typically implemented in object-oriented languages such as Java and Scala due to the quick development cycle they provide. These systems are executed on top of a managed runtime such as the Java Virtual Machine (JVM), which requires each data item to be represented as an object before it can be processed. This representation is the direct cause of many kinds of severe inefficiencies. We developed Gerenuk, a compiler and runtime that aims to enable a JVM-based data-parallel system to achieve near-native efficiency by transforming a set of statements in the system for direct execution over inlined native bytes. The key insight leading to Gerenuk's success is two-fold: (1) analytics workloads often use immutable and confined data types. If we speculatively optimize the system and user code with this assumption, the transformation can be made tractable. (2) The flow of data starts at a deserialization point where objects are created from a sequence of native bytes and ends at a serialization point where they are turned back into a byte sequence to be sent to the disk or network. This flow naturally defines a speculative execution region (SER) to be transformed. Gerenuk compiles a SER speculatively into a version that can operate directly over native bytes that come from the disk or network. The Gerenuk runtime aborts the SER execution upon violations of the immutability and confinement assumption and switches to the slow path by deserializing the bytes and re-executing the original SER. Our evaluation on Spark and Hadoop demonstrates promising results.