Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology

Integration of Large-Scale Data Processing Systems and Traditional Parallel Database Technology
复制标题

DOI:
10.14778/3352063.3352145
复制
发表时间:
2019-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
A. Abouzeid;D. Abadi;Kamil Bajda-Pawlikowski;A. Silberschatz
A. Abouzeid;D. Abadi;Kamil Bajda-Pawlikowski;A. Silberschatz
中科院分区:
其他
文献类型:
--
作者:
A. Abouzeid;D. Abadi;Kamil Bajda-Pawlikowski;A. Silberschatz

文献摘要

相似文献

2009年,我们探索了构建混合SQL数据分析系统的可行性,该系统吸收了两种相互竞争的技术的最佳功能:大型数据处理系统(如Google MapReduced和Apache Hadoop)和并行数据库管理系统(如Greenplum和Vertica)。我们构建了一个原型HadoopDB,并证明了它可以提供并行数据库管理系统的高SQL查询性能和ffi效率,同时仍然提供大型数据处理系统的可伸缩性、容错性和fl灵活性。随后,HadoopDB成长为商业产品HAdapt,其技术最终被Teradata收购。在本文中,我们概述了HadoopDB的原始设计,以及在随后十年的研究和开发过程中ff数据库的演变。我们描述了该项目是如何在研究实验室以及作为HAdapt和Teradata的商业产品进行创新的。然后,我们讨论当前充满活力的软件项目生态系统(其中大部分是开源的),这些项目延续了HadoopDB实现大规模数据处理系统和并行数据库技术的系统级集成的传统。
In 2009 we explored the feasibility of building a hybrid SQL data analysis system that takes the best features from two competing technologies: large-scale data processing systems (such as Google MapReduce and Apache Hadoop) and parallel database management systems (such as Greenplum and Vertica). We built a prototype, HadoopDB, and demonstrated that it can deliver the high SQL query performance and efficiency of parallel database management systems while still providing the scalability, fault tolerance, and flex-ibility of large-scale data processing systems. Subsequently, HadoopDB grew into a commercial product, Hadapt, whose technology was eventually acquired by Teradata. In this paper, we provide an overview of HadoopDB’s original design, and its evolution during the subsequent ten years of research and development effort. We describe how the project inno-vated both in the research lab, and as a commercial product at Hadapt and Teradata. We then discuss the current vibrant ecosystem of software projects (most of which are open source) that continued HadoopDB’s legacy of implementing a systems level integration of large-scale data processing systems and parallel database technology.