Red Fox: An Execution Environment for Relational Query Processing on GPUs

Red Fox: An Execution Environment for Relational Query Processing on GPUs
复制标题

Red Fox:GPU 上关系查询处理的执行环境

DOI:
--
复制
发表时间:
2014
期刊:
IEEE/ACM International Symposium on Code Generation and Optimization
影响因子:
--
通讯作者:
S. Yalamanchili
S. Yalamanchili
中科院分区:
--
文献类型:
--
作者:
Haicheng Wu;G. Diamos;T. Sheard;M. Aref;S. Baxter;M. Garland;S. Yalamanchili

文献摘要

被引文献

相似文献

现代企业应用程序代表了一个新兴的应用程序竞技场,它需要处理大量数据的查询和计算。大规模的多GPU集群系统可能会带来吞吐量和整体性能的重大改进。然而,使用GPU的吞吐量改进受到关系代数(RA)运算符的独特存储器和计算特性的挑战,关系代数(RA)运算符是用于回答业务问题的查询的核心。 本文介绍了Red Fox的设计、实现和评估,Red Fox是一个用于在GPU上执行关系查询的编译器和运行时基础设施。Red Fox由i)LogiQL的语言前端(这是一种商业查询语言),ii)RA到GPU编译器,iii)RA运算符的优化GPU实现,以及iv)支持运行时组成。我们报告了在单节点GPU上执行全套行业标准TPC-H查询的性能。与针对最先进的CPU机器优化的商业LogiQL系统实现相比,Red Fox平均快6.48倍,包括PCIe传输时间。我们指出关键的瓶颈,提出潜在的解决方案,并分析这些查询的GPU实现。据我们所知,这是第一个报告的端到端编译和执行基础设施,支持商品GPU上的全套TPC-H查询。
Modern enterprise applications represent an emergent application arena that requires the processing of queries and computations over massive amounts of data. Large-scale, multi-GPU cluster systems potentially present a vehicle for major improvements in throughput and consequently overall performance. However, throughput improvement using GPUs is challenged by the distinctive memory and computational characteristics of Relational Algebra (RA) operators that are central to queries for answering business questions. This paper introduces the design, implementation, and evaluation of Red Fox, a compiler and runtime infrastructure for executing relational queries on GPUs. Red Fox is comprised of i) a language front-end for LogiQL which is a commercial query language, ii) an RA to GPU compiler, iii) optimized GPU implementation of RA operators, and iv) a supporting runtime. We report the performance on the full set of industry standard TPC-H queries on a single node GPU. Compared with a commercial LogiQL system implementation optimized for a state of art CPU machine, Red Fox on average is 6.48x faster including PCIe transfer time. We point out key bottlenecks, propose potential solutions, and analyze the GPU implementation of these queries. To the best of our knowledge, this is the first reported end-to-end compilation and execution infrastructure that supports the full set of TPC-H queries on commodity GPUs.