The Stratosphere platform for big data analytics

The Stratosphere platform for big data analytics
复制标题

DOI:
10.1007/s00778-014-0357-y
复制
发表时间:
2014-12-01
期刊:
影响因子:
4.2
通讯作者:
Warneke, Daniel
Warneke, Daniel
中科院分区:
计算机科学2区
文献类型:
--
作者:
Alexandrov, Alexander;Bergmann, Rico;Warneke, Daniel

文献摘要

被引文献

相似文献

我们提出了平流层,这是一种用于并行数据分析的开源软件堆栈。平流层汇集了一套独特的功能,这些功能允许在非常大规模的分析应用程序中进行表现力,容易,有效的编程。平流层的功能包括“原位”数据处理,一种声明的查询语言,用户定义的一流公民的处理,自动程序并行化和优化,对迭代程序的支持以及可扩展,有效的执行引擎。平流层涵盖了各种“大数据”用例,例如数据仓库,信息提取和集成,数据清理,图形分析和统计分析应用。在本文中,我们介绍了整个系统体系结构设计决策,通过示例查询介绍平流层,然后深入研究与可扩展性,编程模型,优化和查询执行相关的系统组件的内部工作。我们通过实验将平流层与流行的开源替代方案进行比较,并以未来几年的研究前景结论。
We present Stratosphere, an open-source software stack for parallel data analysis. Stratosphere brings together a unique set of features that allow the expressive, easy, and efficient programming of analytical applications at very large scale. Stratosphere's features include "in situ" data processing, a declarative query language, treatment of user-defined functions as first-class citizens, automatic program parallelization and optimization, support for iterative programs, and a scalable and efficient execution engine. Stratosphere covers a variety of "Big Data" use cases, such as data warehousing, information extraction and integration, data cleansing, graph analysis, and statistical analysis applications. In this paper, we present the overall system architecture design decisions, introduce Stratosphere through example queries, and then dive into the internal workings of the system's components that relate to extensibility, programming model, optimization, and query execution. We experimentally compare Stratosphere against popular open-source alternatives, and we conclude with a research outlook for the next years.