Learning-Based Dynamic Memory Allocation Schemes for Apache Spark Data Processing

Learning-Based Dynamic Memory Allocation Schemes for Apache Spark Data Processing
复制标题

DOI:
10.1109/tcc.2023.3329129
复制
发表时间:
2024-01
影响因子:
6.5
通讯作者:
Danlin Jia;Li Wang;Natalia Valencia;J. Bhimani;Bo Sheng;N. Mi
Danlin Jia;Li Wang;Natalia Valencia;J. Bhimani;Bo Sheng;N. Mi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Danlin Jia;Li Wang;Natalia Valencia;J. Bhimani;Bo Sheng;N. Mi

文献摘要

相似文献

Apache Spark是一个内存分析框架,已被工业和研究领域采用。Spark中有两个内存管理器,Static和Unified,用于分配内存以缓存弹性分布式数据集(RDD)和执行任务。然而,我们发现静态内存管理器(SMM)缺乏灵活性,而统一内存管理器(UMM)给Spark所在的JVM的垃圾收集带来了沉重的压力。为了解决这些问题,我们设计了一个基于学习的双向使用有界的内存分配方案,以支持动态内存分配的内存需求和垃圾收集引入的延迟考虑。我们首先开发了一个自动调整内存管理器(ATuMm),采用了直观的基于反馈的学习解决方案。然而,ATuMm是一个缓慢的学习者,只能在有限的范围内改变Java虚拟内存(JVM)堆的状态。也就是说,ATuMm决定将执行和存储内存池之间的边界增加或减少JVM堆大小的固定部分。为了克服这个缺点,我们进一步开发了一个新的基于强化学习的内存管理器(Q-ATuMm),它使用一个Q-learning智能代理来动态学习和调整JVM堆的分区。我们在Spark 2.4.0中实现了新的内存管理器,并通过在真实的Spark集群中进行实验来评估它们。我们的实验结果表明,我们的内存管理器可以减少总的垃圾收集时间,从而进一步提高Spark应用程序的性能(即,与现有的Spark内存管理解决方案相比,减少了延迟。通过将我们的机器学习驱动的内存管理器集成到Spark中,我们可以进一步将延迟降低约1.3倍。
Apache Spark is an in-memory analytic framework that has been adopted in the industry and research fields. Two memory managers, Static and Unified, are available in Spark to allocate memory for caching Resilient Distributed Datasets (RDDs) and executing tasks. However, we find that the static memory manager (SMM) lacks flexibility, while the unified memory manager (UMM) puts heavy pressure on the garbage collection of the JVM on which Spark resides. To address these issues, we design a learning-based bidirectional usage-bounded memory allocation scheme to support dynamic memory allocation with the consideration of both memory demands and latency introduced by garbage collection. We first develop an auto-tuning memory manager (ATuMm) that adopts an intuitive feedback-based learning solution. However, ATuMm is a slow learner that can only alter the states of Java Virtual Memory (JVM) Heap in a limited range. That is, ATuMm decides to increase or decrease the boundary between the execution and storage memory pools by a fixed portion of JVM Heap size. To overcome this shortcoming, we further develop a new reinforcement learning-based memory manager (Q-ATuMm) that uses a Q-learning intelligent agent to dynamically learn and tune the partition of JVM Heap. We implement our new memory managers in Spark 2.4.0 and evaluate them by conducting experiments in a real Spark cluster. Our experimental results show that our memory manager can reduce the total garbage collection time and thus further improve Spark applications’ performance (i.e., reduced latency) compared to the existing Spark memory management solutions. By integrating our machine learning-driven memory manager into Spark, we can further obtain around 1.3x times reduction in the latency.