SMURF: Efficient and Scalable Metadata Access for Distributed Applications

SMURF: Efficient and Scalable Metadata Access for Distributed Applications
复制标题

DOI:
10.1109/tpds.2022.3175596
复制
发表时间:
2021-05
影响因子:
5.3
通讯作者:
Bing Zhang;T. Kosar
Bing Zhang;T. Kosar
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bing Zhang;T. Kosar

文献摘要

被引文献

相似文献

随着大数据处理和分析主导分布式和云基础设施的使用,对分布式元数据访问和传输的需求也在增加。许多应用程序域生成的数据量超过pb,而相应的元数据达到tb甚至更多。本文提出了一种新的解决方案,用于跨广域网的分布式应用程序的高效和可扩展的元数据访问,称为SMURF。我们的解决方案将新颖的流水线和并发传输机制与可靠性相结合,提供分布式连续缓存和语义位置感知预取策略,以避免获取延迟,并在云中实现可扩展和高性能的元数据获取/预取服务。我们结合了语义位置感知现象来提高预取预测率,使用来自Yahoo!Hadoop审计日志并提出一种新的预取预测器。通过基于访问模式有效地缓存和预取元数据,我们的连续体缓存和预取机制显著提高了本地缓存命中率,降低了平均提取延迟。我们从真实的审计跟踪中重播了大约2000万次元数据访问操作,其中SMURF在预取预测期间达到了90%的准确率,与最先进的机制相比,平均提取延迟减少了50%。
In parallel with big data processing and analysis dominating the usage of distributed and Cloud infrastructures, the demand for distributed metadata access and transfer has increased. The volume of data generated by many application domains exceeds petabytes, while the corresponding metadata amounts to terabytes or even more. This paper proposes a novel solution for efficient and scalable metadata access for distributed applications across wide-area networks, dubbed SMURF. Our solution combines novel pipelining and concurrent transfer mechanisms with reliability, provides distributed continuum caching and semantic locality-aware prefetching strategies to sidestep fetching latency, and achieves scalable and high-performance metadata fetch/prefetch services in the Cloud. We incorporate the phenomenon of semantic locality awareness for increased prefetch prediction rate using real-life application I/O traces from Yahoo! Hadoop audit logs and propose a novel prefetch predictor. By effectively caching and prefetching metadata based on the access patterns, our continuum caching and prefetching mechanism significantly improves the local cache hit rate and reduces the average fetching latency. We replay approximately 20 Million metadata access operations from real audit traces, where SMURF achieves 90% accuracy during prefetch prediction and reduced the average fetch latency by 50% compared to the state-of-the-art mechanisms.