课题基金 / 基金详情

A Scalable, Massively-Parallel Runtime System with Predictable Performance

A Scalable, Massively-Parallel Runtime System with Predictable Performance
具有可预测性能的可扩展、大规模并行运行时系统
批准号:
248358398
负责人:
Professor Dr. Odej Kao
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Units
财政年份:
2013
资助国家:
德国
项目状态:
已结题
起止时间:
2012-12-31 至 2016-12-31

项目摘要

项目成果

Professor Dr. Odej Kao的其他基金

相似基金

相关文献

中文摘要
翻译
Stratosphere II研究部内项目B的目标是研究和开发运行时系统“Aura”,以便在分布式云或集群基础设施上以大规模并行方式执行多个并发数据分析程序。项目B分为两个主要研究领域RA-B.1和RA-B.2,它们增强了运行时环境,使其具有处理不断发展的数据集以及表示和管理所产生的分布式状态的能力,目标是以大规模并行、容错的方式支持新颖的迭代数据分析算法。我们计划建立一个新的执行模型,即所谓的Supremo执行计划(SEP),它以执行的形式对框架的工作负载进行建模。SEP是一个受限的循环图,它以物理视图的形式将不断发展的数据集与语义丰富的运算符相结合。物理视图可以表示传统的数据源和数据汇,但也能够在迭代数据分析中保持状态,以及在无限数据上的有状态运算符中出现的状态,例如窗口运算符。与Stratosphere I的Nephele的UDF黑盒相比,关于操作员特征的附加知识与工作负载感知(重新)调度策略相结合,允许运行时核心的调度程序提供单个部署的数据分析程序的可预测运行时行为。具体而言,本项目旨在回答以下问题:1。如何构建运行时系统,以优化各种硬件架构上的迭代数据分析程序的执行,利用虚拟化硬件的优势?2.我们如何在大型计算集群上有效地维护状态并提供具有迭代的程序的容错执行?3.我们如何适应虚拟化方法的特点,在低延迟边界和资源保证方面实现可预测的性能?假设云系统具有按需弹性,我们如何根据计算需求或摄取率提供向上和向下扩展?4.如何在并发查询的工作负载之间共享和分布大型复杂模型?
英文摘要
The goal of project B within the Stratosphere II Research Unit is to research and develop the runtime system “Aura” to execute multiple, concurrent data analysis programs in a massively parallel fashion on a distributed cloud or cluster infrastructure. Project B is divided into two primary research areas RA-B.1 and RA-B.2 that enhance the runtime environment with capabilities to handle evolving datasets and to represent and manage the resulting distributed state with the goal of supporting novel iterative data analysis algorithms in a massively parallel, fault-tolerant way. We plan to establish a novel execution model, the so-called Supremo Execution Plan (SEP), which models the framework’s workload in the form in which it is executed. A SEP is a restricted cyclic graph that combines evolving datasets in the form of physical views, with semantically rich operators. Physical views can represent traditional data sources and sinks but will also be able to hold the state in iterative data analysis, as well as the state occurring in stateful operators on infinite data, e.g. windowed operators. In contrast to the UDF black-boxes of Stratosphere I’s Nephele, the added knowledge about operator characteristics in combination with workload-aware (re-)scheduling policies allows the runtime core’s scheduler to provide predictable runtime behavior of individual deployed data analysis programs. In particular, this project aims at answering following questions:1. How must a runtime system be architected to optimize for the execution of iterative data analysis programs on various hardware architectures, exploiting the advantages of a virtualized hardware?2. How can we efficiently maintain state and provide fault-tolerant execution of programs with iterations on large-compute clusters?3. How can we adapt to the characteristics of virtualization methods to achieve predictable performance in terms of low-latency bounds and resource guarantees? How do we provide up- and down-scaling based on computational needs or on ingestion rates, assuming on-demand elasticity of Cloud systems?4. How can large, complex models be shared and distributed between workloads of concurrent queries?
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/cse.2013.186
发表时间: 2013
期刊: 2013 IEEE 16th International Conference on Computational Science and Engineering
影响因子: --
作者: [Mareike Höger, Odej Kao]
通讯作者: Odej Kao
DOI: 10.1109/bigdata.2015.7364083
发表时间: 2015-10
期刊: 2015 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者: [T. Renner;L. Thamsen;O. Kao]
通讯作者: T. Renner;L. Thamsen;O. Kao
DOI: 10.1109/cloud.2011.30
发表时间: 2011-07
期刊: 2011 IEEE 4th International Conference on Cloud Computing
影响因子: --
作者: [Dominic Battré;Natalia Frejnik;Siddhant Goel;O. Kao;Daniel Warneke]
通讯作者: Dominic Battré;Natalia Frejnik;Siddhant Goel;O. Kao;Daniel Warneke
Inferring Network Topologies in Infrastructure as a Service Cloud
推断基础设施即服务云中的网络拓扑
DOI: 10.1109/ccgrid.2011.79
发表时间: 2011
期刊: 2011 11th IEEE/ACM International Symposium on Cluster, Cloud and Grid Computing
影响因子: --
作者: [Dominic Battré, Natalia Frejnik, Siddhant Goel, Odej Kao, Daniel Warneke]
通讯作者: Daniel Warneke
共 8 条
    Massively Parallel, Adaptive and Fault-Tolerant Execution of Data Flow Programs on Dynamic Clouds
    C5: Collaborative and Cross-Context Cluster Configuration for Distributed Data-Parallel Processing
    海外基金