Effortless Locality on Data Systems Using Relational Fabric

Effortless Locality on Data Systems Using Relational Fabric
复制标题

DOI:
10.1109/tkde.2024.3386827
复制
发表时间:
2024
影响因子:
8.9
通讯作者:
Tarikul Islam Papon;J. Mun;Konstantinos Karatsenidis;Shahin Roozkhosh;Denis Hoornaert;Ahmed Sanaullah;Ulrich Drepper;Renato Mancuso;Manos Athanassoulis
Tarikul Islam Papon;J. Mun;Konstantinos Karatsenidis;Shahin Roozkhosh;Denis Hoornaert;Ahmed Sanaullah;Ulrich Drepper;Renato Mancuso;Manos Athanassoulis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tarikul Islam Papon;J. Mun;Konstantinos Karatsenidis;Shahin Roozkhosh;Denis Hoornaert;Ahmed Sanaullah;Ulrich Drepper;Renato Mancuso;Manos Athanassoulis

文献摘要

相似文献

数据系统的一个关键设计决策是它们是否遵循行存储或列存储范式。前者支持事务工作负载,而后者更适合分析查询。这一决定对整个数据系统架构有重大影响。这两种设计的几十年之久的旅程导致了一个新的混合事务/分析处理(HTAP)架构家族。已经提出了一些努力,通过提出维护多个数据副本(在不同的物理布局中)并根据需要将它们转换为所需布局的系统来获得两个世界的好处。由于数据重复,额外的必要簿记以及在不同布局之间转换数据的成本,这些系统在有效分析和数据新鲜度之间进行了妥协。我们从现有的设计出发,提出了一种全新的方法。我们提出这样一个问题:“如果我们可以访问任何布局,并通过透明地将行转换为(任意组的)列,通过内存层次结构只传输相关数据,那会怎么样?”为了实现这一功能,我们利用硬件专业化的复兴趋势(由于摩尔定律的逐渐减弱而加速),提出了关系结构,一种接近数据的垂直分区器,允许内存或存储组件执行动态透明数据转换。通过公开一个直观的API,Relational Fabric将垂直分区推向硬件,这对设计和构建数据系统的过程产生了深远的影响。(A)不需要数据复制和布局转换,使HTAP系统可以使用单一布局。(B)它简化了需要维护和更新单个数据布局的内存和存储管理器。(C)它减少了内存层次结构中不必要的数据移动,从而提高了硬件利用率,并最终提高了性能。在本文中,我们提出了关系结构的内存和存储。我们提出了我们的关系结构的内存系统的初步结果,并讨论了构建这种硬件的挑战和机遇,它带来的简单性和创新的数据系统软件栈,包括物理设计,查询优化,查询评估和并发控制。
—A key design decision for data systems is whether they follow the row-store or the column-store paradigm. The former supports transactional workloads, while the latter is better for analytical queries. This decision has a significant impact on the entire data system architecture. The multiple-decade-long journey of these two designs has led to a new family of hybrid transactional/analytical processing (HTAP) architectures. Several efforts have been proposed to reap the benefits of both worlds by proposing systems that maintain multiple copies of data (in different physical layouts) and convert them into the desired layout as required. Due to data duplication, the additional necessary bookkeeping, and the cost of converting data between different layouts, these systems compromise between efficient analytics and data freshness. We depart from existing designs by proposing a radically new approach. We ask the question: “What if we could access any layout and ship only the relevant data through the memory hierarchy by transparently converting rows to (arbitrary groups of) columns?” To achieve this functionality, we capitalize on the reinvigorated trend of hardware specialization (that has been accelerated due to the tapering of Moore’s law) to propose Relational Fabric , a near-data vertical partitioner that allows memory or storage components to perform on-the-fly transparent data transformation. By exposing an intuitive API, Relational Fabric pushes vertical partitioning to the hardware, which profoundly impacts the process of designing and building data systems. (A) There is no need for data duplication and layout conversion, making HTAP systems viable using a single layout. (B) It simplifies the memory and storage manager that needs to maintain and update a single data layout. (C) It reduces unnecessary data movement through the memory hierarchy allowing for better hardware utilization and, ultimately, better performance. In this paper, we present Relational Fabric for both memory and storage. We present our initial results on Relational Fabric for in-memory systems and discuss the challenges of building this hardware and the opportunities it brings for simplicity and innovation in the data system software stack, including physical design, query optimization, query evaluation, and concurrency control.