Improving Apache Spark's Cache Mechanism with LRC-Based Method Using Bloom Filter

Improving Apache Spark's Cache Mechanism with LRC-Based Method Using Bloom Filter
复制标题

DOI:
10.1109/candarw.2018.00096
复制
发表时间:
2018-11
期刊:
2018 Sixth International Symposium on Computing and Networking Workshops (CANDARW)
影响因子:
--
通讯作者:
Hideo Inagaki;Ryota Kawashima;H. Matsuo
Hideo Inagaki;Ryota Kawashima;H. Matsuo
中科院分区:
其他
文献类型:
--
作者:
Hideo Inagaki;Ryota Kawashima;H. Matsuo

文献摘要

相似文献

内存和磁盘缓存是Apache Spark中时间输出的常见缓存机制。但是,由于基于Spark的LRU(基于最不最近使用的)高速缓存管理,这会导致性能下降。现有研究报告说,基于LRU的缓存机制替换为LRC(最少参考计数),这是更准确的未来数据访问可能性的指标。但是,无法确定经常使用的分区,因为Spark可以访问用户驱动的RDD操作的所有分区,即使分区不包含必要的数据。在本文中,我们提出了一种缓存管理方法,该方法可以通过将Bloom过滤器引入现有方法来将必要的分区分配给内存。 Bloom过滤器可防止不必要的分区进行处理,因为检查了是否包含所需数据。此外,可以通过测量分区的参考计数来正确确定经常使用的分区。我们实现了两种体系结构类型,即驱动程序侧花滤波器和执行者侧花过滤器,以考虑Bloom滤波器的最佳位置。评估结果表明,基于基于LRC的方法的滤镜测试基准,驾驶员端实现的执行时间减少了89%。
Memory-and-Disk caching is a common caching mechanism for temporal output in Apache Spark. However, it causes performance degradation when memory usage has reached its limit because of the Spark's LRU (Least Recently Used) based cache management. Existing studies have reported that replacement of LRU-based cache mechanism to LRC (Least Reference Count) based one that is a more accurate indicator of the likelihood of future data access. However, frequently used partitions cannot be determined because Spark accesses all of partitions for user-driven RDD operations, even if partitions do not include necessary data. In this paper, we propose a cache management method that enables allocating necessary partitions to the memory by introducing the bloom filter into existing methods. The bloom filter prevents unnecessary partitions from being processed because partitions are checked whether required data is contained. Furthermore, frequently used partitions can be properly determined by measuring the reference count of partitions. We implemented two architecture types, the driver-side bloom filter and the executor-side bloom filter, to consider the optimal place of the bloom filter. Evaluation results showed that the execution time of the driver-side implementation was reduced by 89% in a filter-test benchmark based on the LRC-based method.