Fast Parallel Recovery of Many Small In-Memory Objects

Fast Parallel Recovery of Many Small In-Memory Objects
复制标题

DOI:
10.1109/icpads.2017.00042
复制
发表时间:
2017-12
期刊:
2017 IEEE 23rd International Conference on Parallel and Distributed Systems (ICPADS)
影响因子:
--
通讯作者:
Kevin Beineke;Stefan Nothaas;M. Schöttner
Kevin Beineke;Stefan Nothaas;M. Schöttner
中科院分区:
其他
文献类型:
--
作者:
Kevin Beineke;Stefan Nothaas;M. Schöttner

文献摘要

被引文献

相似文献

社交媒体网络以及在线图分析在具有数百万个顶点的大规模图上运行,在某些情况下甚至是数十亿个顶点。低延迟访问是必不可少的,但缓存受到上述应用程序域的大多数不规则访问模式的影响。因此,提出了分布式内存系统,将所有数据始终保持在内存中。这些内存系统通常没有针对大量的小数据对象进行优化,这需要关于本地和全局数据管理的新概念以及用于掩盖服务器故障和停电的容错机制。在本文中,我们提出了一种新的备份分布和并行恢复方法,旨在快速恢复存储数以亿计的小对象的服务器。所有提出的概念都已在开源分布式系统DXRAM中实现,并在Microsoft Azure云中进行了评估,在两个规模集中有多达72个高性能虚拟机。为了进行评估,我们使用了两个基准测试:Yahoo!云服务基准和恢复基准。实验表明,该恢复策略能够恢复服务器与5亿小数据对象在不到2秒,也,有效?在重负载下屏蔽服务器故障。此外,DXRAM在大型对象的额外恢复实验中表现优于最先进的系统RAMCloud(速度快2.4倍),在小型对象的恢复实验中表现更好(> 9倍)。
Social media networks as well as online graph analytics operate on large-scale graphs with millions of vertices, even billions in some cases. Low-latency access is essential, but caching suffers from the mostly irregular access patterns of the aforementioned application domains. Hence, distributed in-memory systems are proposed keeping all data always in memory. These in-memory systems are typically not optimized for the sheer amounts of small data objects, which demands new concepts regarding the local and global data management as well as for the fault-tolerance mechanisms to mask server failures and power outages. In this paper, we propose a novel backup distribution and parallel recovery approach aiming at fast recovery of servers storing hundreds of millions of small objects. All proposed concepts have been implemented within the open source distributed system DXRAM and have been evaluated in the Microsoft Azure cloud with up to 72 high performance virtual machines in two scale-sets. For evaluation, we used two benchmarks: the Yahoo! Cloud Serving Benchmark and a recovery benchmark. The experiments show that the proposed recovery strategy is able to recover servers with 500,000,000 small data objects in less than 2 seconds and, also, to ef?ciently mask server failures under heavy load. Furthermore, DXRAM outperforms the state-of-the-art system RAMCloud in additional recovery experiments with large objects (2.4x faster) and even more with small objects (> 9x).