Robust and efficient memory management in Apache AsterixDB

Robust and efficient memory management in Apache AsterixDB
复制标题

DOI:
10.1002/spe.2799
复制
发表时间:
2020-02
期刊:
Software: Practice and Experience
影响因子:
--
通讯作者:
Taewoo Kim;Alexander Behm;Michael Blow;V. Borkar;Yingyi Bu;M. Carey;Murtadha Ai Hubail;Shiva Jahangiri;Jianfeng Jia;Chen Li;Chen Luo;Ian Maxon;Pouria Pirzadeh
Taewoo Kim;Alexander Behm;Michael Blow;V. Borkar;Yingyi Bu;M. Carey;Murtadha Ai Hubail;Shiva Jahangiri;Jianfeng Jia;Chen Li;Chen Luo;Ian Maxon;Pouria Pirzadeh
中科院分区:
其他
文献类型:
--
作者:
Taewoo Kim;Alexander Behm;Michael Blow;V. Borkar;Yingyi Bu;M. Carey;Murtadha Ai Hubail;Shiva Jahangiri;Jianfeng Jia;Chen Li;Chen Luo;Ian Maxon;Pouria Pirzadeh

文献摘要

被引文献

相似文献

传统的关系数据库系统通过将其内存划分为诸如缓冲器高速缓存和工作内存等部分来处理数据,并为每个部分分配内存预算以有效地管理有限数量的总内存。它们还将内存预算分配给排序和联接等内存密集型操作符,并控制这些操作符的内存分配;每个内存密集型操作符都试图最大化其内存使用量,以降低磁盘I/O成本。实现这种内存密集型运算符需要仔细设计和应用适当的算法来正确利用内存。今天的大数据管理系统需要以类似的方式处理大量数据的能力,因为假设真正的大数据将适合内存是不现实的。在本文中,我们将分享我们在ApacheAsterixDB中的内存管理经验,这是一个开源大数据管理软件平台,在无共享的商用计算集群上横向扩展。我们描述了AsterixDB的内存密集型运算符的实现及其与内存管理相关的设计。我们还讨论了全局(集群)级别的内存管理。我们使用几个合成和真实的数据集进行了一项实验研究,以探索这项工作的影响。我们相信,未来的大数据管理系统建设者可以从这些经验中受益。
Traditional relational database systems handle data by dividing their memory into sections such as a buffer cache and working memory, assigning a memory budget to each section to efficiently manage a limited amount of overall memory. They also assign memory budgets to memory‐intensive operators such as sorts and joins and control the allocation of memory to these operators; each memory‐intensive operator attempts to maximize its memory usage to reduce disk I/O cost. Implementing such memory‐intensive operators requires a careful design and application of appropriate algorithms that properly utilize memory. Today's Big Data management systems need the ability to handle large amounts of data similarly, as it is unrealistic to assume that truly big data will fit into memory. In this article, we share our memory management experiences in Apache AsterixDB, an open‐source Big Data management software platform that scales out horizontally on shared‐nothing commodity computing clusters. We describe the implementation of AsterixDB's memory‐intensive operators and their designs related to memory management. We also discuss memory management at the global (cluster) level. We conducted an experimental study using several synthetic and real datasets to explore the impact of this work. We believe that future Big Data management system builders can benefit from these experiences.