Exploration of memory hybridization for RDD caching in Spark

Exploration of memory hybridization for RDD caching in Spark
复制标题

DOI:
10.1145/3315573.3329988
复制
发表时间:
2019-06
期刊:
Proceedings of the 2019 ACM SIGPLAN International Symposium on Memory Management
影响因子:
--
通讯作者:
Md. Muhib Khan;Muhammad Ahad Ul Alam;Amit Kumar Nath;Weikuan Yu
Md. Muhib Khan;Muhammad Ahad Ul Alam;Amit Kumar Nath;Weikuan Yu
中科院分区:
其他
文献类型:
--
作者:
Md. Muhib Khan;Muhammad Ahad Ul Alam;Amit Kumar Nath;Weikuan Yu

文献摘要

相似文献

由于使用弹性分布式数据集(RDDS)来缓存数据以进行内存处理,因此Apache Spark是适用于迭代分析工作负载的流行集群计算框架。我们已经透露,如果Spark RDD缓存的容量不能满足工作负载的需求,其性能可能会受到严重限制。在本文中,我们探索了不同的内存混合策略,以将紧急非易失性内存(NVM)设备用于Spark的RDD缓存。我们发现,简单的分层杂交方法不能提供有效的解决方案。因此,我们设计了一种平面混合方案来利用NVM来缓存RDD块,以及几种架构优化,如用于块展开的动态内存分配、带抢占的异步迁移和机会性磁盘逐出。我们已经进行了一系列广泛的实验来评估我们提出的平面杂交策略的性能,并发现它在处理不同的系统和NVM特征时是稳健的。我们提出的方法将DRAM用于混合存储系统的一小部分,但仍设法将执行时间的增加平均控制在10%以内。此外,当与当前机制配合使用时,我们的机会性数据块回收到磁盘可将性能提高高达7.5%。
Apache Spark is a popular cluster computing framework for iterative analytics workloads due to its use of Resilient Distributed Datasets (RDDs) to cache data for in-memory processing. We have revealed that the performance of Spark RDD cache can be severely limited if its capacity falls short to the needs of the workloads. In this paper, we have explored different memory hybridization strategies to leverage emergent Non-Volatile Memory (NVM) devices for Spark's RDD cache. We have found that a simple layered hybridization approach does not offer an effective solution. Therefore, we have designed a flat hybridization scheme to leverage NVM for caching RDD blocks, along with several architectural optimizations such as dynamic memory allocation for block unrolling, asynchronous migration with preemption, and opportunistic eviction to disk. We have performed an extensive set of experiments to evaluate the performance of our proposed flat hybridization strategy and found it to be robust in handling different system and NVM characteristics. Our proposed approach uses DRAM for a fraction of the hybrid memory system and yet manages to keep the increase in execution time to be within 10% on average. Moreover, our opportunistic eviction of blocks to disk improves performance by up to 7.5% when utilized alongside the current mechanism.