Understanding and improving disk-based intermediate data caching in Spark

Understanding and improving disk-based intermediate data caching in Spark
复制标题

DOI:
10.1109/bigdata.2017.8258209
复制
发表时间:
2017-12
期刊:
2017 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Kaihui Zhang;Y. Tanimura;H. Nakada;Hirotaka Ogawa
Kaihui Zhang;Y. Tanimura;H. Nakada;Hirotaka Ogawa
中科院分区:
其他
文献类型:
--
作者:
Kaihui Zhang;Y. Tanimura;H. Nakada;Hirotaka Ogawa

文献摘要

相似文献

Apache Spark是一个并行数据处理框架,通过在内存中缓存中间数据,并在故障时进行基于沿袭的数据恢复,可以快速执行迭代计算和交互式处理。Spark系统还可以通过在处理节点的磁盘上放置部分缓存或全部缓存来管理大于内存容量的数据集。然而,缺点是由于磁盘I/O和/或序列化而导致的潜在性能下降。这项研究的目的是澄清在Spark中间数据缓存中磁盘的高效/低效使用,并提高磁盘对最终用户的可用性。为了达到这一目的,首先从缓存选项、数据抽象和存储设备等方面研究了磁盘使用对数据缓存的影响。结果表明,在大多数情况下,序列化成本占主导地位,而不是磁盘I/O成本。其次,在内存压力较大的情况下,进一步评估了一种内存与磁盘结合使用的方法。然后对该方法进行了改进,以避免过度的重新缓存问题,在高内存压力下,该方法最多减少了20-30%的总执行时间,并且在低内存压力下,我们对4个机器学习基准进行了实验。最后,本文总结了在Spark中有效使用磁盘进行数据缓存的重要因素和潜在改进。
Apache Spark is a parallel data processing framework that executes fast for iterative calculations and interactive processing, by caching intermediate data in memory with a lineage-based data recovery from faults. The Spark system can also manage data sets larger than memory capacity by placing some cache or all of them on disks on processing nodes. However, the disadvantage is potential performance degradation due to disk I/O and/or serialization. This study aims to clarify efficient/inefficient use of disks in intermediate data caching in Spark and also to improve the usability of disks for end users. In order to achieve the purpose, influence of disk use in data caching was firstly investigated in various aspects, such as caching options, data abstractions and storage devices. The results indicate that serialization cost is dominant rather than disk I/O in most cases. Secondly, a method of combined use of memory and disk was further evaluated under a high memory pressure. Then the method was improved to avoid an excessive re-caching problem, which achieved at most 20–30% reduction of total execution time under a high memory pressure and did not degrade the performance under a low memory pressure, in our experiment with 4 machine learning benchmarks. Finally, this paper summarizes important factors and potential improvements for efficiently using disks in data caching in Spark.