HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads

HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads
复制标题

DOI:
10.14778/1687627.1687731
复制
发表时间:
2009-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
A. Abouzeid;Kamil Bajda-Pawlikowski;D. Abadi;A. Rasin;A. Silberschatz
A. Abouzeid;Kamil Bajda-Pawlikowski;D. Abadi;A. Rasin;A. Silberschatz
中科院分区:
其他
文献类型:
--
作者:
A. Abouzeid;Kamil Bajda-Pawlikowski;D. Abadi;A. Rasin;A. Silberschatz

文献摘要

被引文献

相似文献

分析数据管理应用的生产环境正在迅速变化。许多企业正在从在高端专有机器上部署其分析数据库转向更便宜、低端的商用硬件,这些硬件通常采用无共享的大规模并行处理(MPP)架构,且常常处于公共或私有“云”内的虚拟化环境中。与此同时,需要分析的数据量呈爆炸式增长,需要数百到数千台机器并行工作来执行分析。在这种环境下使用何种技术进行数据分析,往往存在两种观点。并行数据库的支持者认为,并行数据库对性能和效率的高度重视使其非常适合执行此类分析。另一方面,其他人则认为基于MapReduce的系统更合适,因为它们具有出色的可扩展性、容错性以及处理非结构化数据的灵活性。在本文中,我们探讨了构建一种混合系统的可行性,该系统融合了两种技术的最佳特性;我们构建的原型在性能和效率上接近并行数据库,但仍然具备基于MapReduce的系统的可扩展性、容错性和灵活性。
The production environment for analytical data management applications is rapidly changing. Many enterprises are shifting away from deploying their analytical databases on high-end proprietary machines, and moving towards cheaper, lower-end, commodity hardware, typically arranged in a shared-nothing MPP architecture, often in a virtualized environment inside public or private "clouds". At the same time, the amount of data that needs to be analyzed is exploding, requiring hundreds to thousands of machines to work in parallel to perform the analysis. There tend to be two schools of thought regarding what technology to use for data analysis in such an environment. Proponents of parallel databases argue that the strong emphasis on performance and efficiency of parallel databases makes them well-suited to perform such analysis. On the other hand, others argue that MapReduce-based systems are better suited due to their superior scalability, fault tolerance, and flexibility to handle unstructured data. In this paper, we explore the feasibility of building a hybrid system that takes the best features from both technologies; the prototype we built approaches parallel databases in performance and efficiency, yet still yields the scalability, fault tolerance, and flexibility of MapReduce-based systems.