On inter-operator data transfers in query processing

On inter-operator data transfers in query processing
复制标题

DOI:
10.1109/icde53745.2022.00066
复制
发表时间:
2022-05
期刊:
2022 IEEE 38th International Conference on Data Engineering (ICDE)
影响因子:
--
通讯作者:
Harshad Deshmukh;Bruhathi Sundarmurthy;Jignesh M. Patel
Harshad Deshmukh;Bruhathi Sundarmurthy;Jignesh M. Patel
中科院分区:
其他
文献类型:
--
作者:
Harshad Deshmukh;Bruhathi Sundarmurthy;Jignesh M. Patel

文献摘要

相似文献

在设计查询处理原语时,一个关键的设计选择是查询计划中两个运算符之间的数据传输方法。当我们考虑我们正在构建的内存数据库系统的关键设计机制时,我们很快意识到(令人惊讶的是)这个概念没有明确的定义。论文中充满了诸如管道和阻塞等术语的临时使用,但这些术语的定义并不明确,因此很难完全理解这些概念所带来的结果。为了解决这个限制,我们引入了一个明确的术语来说明如何考虑查询管道中运算符之间的数据传输。我们认为,管道和阻塞没有明确的定义,并且存在基于称为传输单元的简单概念的全套技术。接下来,我们开发一个用于操作员间通信的分析模型,并突出显示影响性能的关键参数(对于内存数据库设置)。有了这个模型,我们就可以将其应用到我们正在设计的系统中,并强调我们从这次练习中收集到的见解。我们发现传统的“流水线”和“非流水线”查询处理方法之间的差距。性能和内存占用等关键因素非常有限,因此系统设计人员可能应该重新考虑内存数据库系统的“流水线”与“阻塞”的概念。
In designing query processing primitives, a crucial design choice is the method for data transfer between two operators in a query plan. As we were considering this critical design mechanism for an in-memory database system that we are building, we quickly realized that (surprisingly) there isn't a clear definition of this concept. Papers are full of ad hoc use of terms like pipelining and blocking, but these terms are not crisply defined, making it hard to fully understand the results attributed to these concepts. To address this limitation, we introduce a clear terminology for how to think about data transfer between operators in a query pipeline. We argue that there isn't a clear definition of pipelining and blocking, and that there is a full spectrum of techniques based on a simple concept called unit-of-transfer. Next, we develop an analytical model for inter-operator communication, and highlight the key parameters that impact performance (for in-memory database settings). Armed with this model, we then apply it to the system we are designing and highlight the insights that we gathered from this exercise. We find that the gap between the traditional “pipelining” and “non-pipelining” methods of query processing, w.r.t. key factors such as performance and memory footprint is quite narrow, and thus system designers should likely rethink the notion of “pipelining” vs. “blocking” for in-memory database systems.