Relational Memory: Native In-Memory Accesses on Rows and Columns

Relational Memory: Native In-Memory Accesses on Rows and Columns
复制标题

DOI:
10.48786/edbt.2023.06
复制
发表时间:
2021-09
影响因子:
--
通讯作者:
Shahin Roozkhosh;Denis Hoornaert;J. Mun;Tarikul Islam Papon;Ulrich Drepper;R. Mancuso;Manos Athanassoulis
Shahin Roozkhosh;Denis Hoornaert;J. Mun;Tarikul Islam Papon;Ulrich Drepper;R. Mancuso;Manos Athanassoulis
中科院分区:
--
文献类型:
--
作者:
Shahin Roozkhosh;Denis Hoornaert;J. Mun;Tarikul Islam Papon;Ulrich Drepper;R. Mancuso;Manos Athanassoulis

文献摘要

被引文献

相似文献

分析数据库系统通常被设计为使用列优先数据布局仅访问所需字段。另一方面,存储数据行首先非常适合访问,插入或更新整个行。在运行时将行转换为列是昂贵的,因此,许多分析系统以排名第一的形式摄入数据,并在背景中将其转换为列以促进未来的分析查询。如果我们始终只能有效访问所需的一组列,该设计将如何更改?为了解决这个问题,我们提出了一种从行到列的数据转换的根本新方法。我们建立在具有可重新编程逻辑的嵌入式平台中的最新进步基础上,以在行和列上设计本机内存访问。我们的方法称为关系内存,依赖于基于FPGA的加速器,该加速器位于CPU和主内存之间,并透明地将基本数据转换为运行时最小开销的任何列组。此设计允许访问任何列组,就好像它已经存在于内存中一样。我们在真实硬件中实现和部署关系内存,我们表明我们可以比从行列方面访问它们的访问速度快1.63倍,同时匹配纯柱子的性能以获得低投影率,并且超越了投影率,并且表现优于表现。随着投影率(和元组再建造成本)的增加,它最多可达1.87倍。此外,我们的方法很容易扩展,以支持将许多操作卸载到硬件上,例如选择,组合,聚合和加入,有可能极大地简化软件逻辑并加速查询执行。
Analytical database systems are typically designed to use a column-first data layout to access only the desired fields. On the other hand, storing data row-first works great for accessing, inserting, or updating entire rows. Transforming rows to columns at runtime is expensive, hence, many analytical systems ingest data in row-first form and transform it in the background to columns to facilitate future analytical queries. How will this design change if we can always efficiently access only the desired set of columns? To address this question, we present a radically new approach to data transformation from rows to columns. We build upon recent advancements in embedded platforms with re-programmable logic to design native in-memory access on rows and columns. Our approach, termed Relational Memory, relies on an FPGA- based accelerator that sits between the CPU and main memory and transparently transforms base data to any group of columns with minimal overhead at runtime. This design allows accessing any group of columns as if it already exists in memory. We implement and deploy Relational Memory in real hardware, and we show that we can access the desired columns up to 1.63x faster than accessing them from their row-wise counterpart, while matching the performance of a pure columnar access for low projectivity, and outperforming it by up to 1.87x as projectivity (and tuple re-construction cost) increases. Moreover, our approach can be easily extended to support offloading of a number of operations to hardware, e.g., selection, group by, aggregation, and joins, having the potential to vastly simplify the software logic and accelerate the query execution.