课题基金 / 基金详情

C5: Collaborative and Cross-Context Cluster Configuration for Distributed Data-Parallel Processing

C5: Collaborative and Cross-Context Cluster Configuration for Distributed Data-Parallel Processing
C5:分布式数据并行处理的协作和跨上下文集群配置
批准号:
506529034
负责人:
Professor Dr. Odej Kao
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Professor Dr. Odej Kao的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Many organizations routinely analyze large datasets today. For this, they make use of distributed data-parallel processing systems and take advantage of clusters of commodity resources. Especially smaller organizations and individual users are enabled by data processing frameworks and cloud computing, allowing them to work with large datasets at a high-level of abstraction. Still, users are required to configure adequate resources for their data processing jobs. This is often not straightforward and users frequently overprovision resources for their jobs, leading to low resource utilization as well as high costs and energy consumptions. Numerous works addressed this problem in the last decade for big data frameworks, scientific workflows, and machine learning systems, using statistical tools and performance models. However, much of the effort focused on industry settings, either assuming data on previous executions of jobs to be available or relying on potentially costly dedicated profiling. Little research has addressed use cases where runtime data is not as easily available. Addressing this research gap, we aim to develop new methods for the collaborative usage of runtime data in the proposed project, C5. We believe sharing of runtime information across different execution contexts presents a significant opportunity for performance modeling and model-based resource management in many situations, especially when the availability of runtime data is limited, and will improve the efficiency of distributed data-parallel processing. The methods we plan to develop and evaluate in this project include: - Similarity measures for computational resources and processing jobs to support the use of runtime data and performance models across execution contexts - Model selection and combination methods for robust performance estimations, even if limited training data is available or model components were trained in other contexts - Adjustment strategies that allow to efficiently update training data, performance models, and resource configurations at runtime. In addition to new methods for cross-context cluster configuration optimization based on shared performance data and models, we plan to conduct a thorough analysis of real workloads, design reproducible experiments based on infrastructure-as-code definitions and benchmarks, and provide a working implementation of the envisioned data sharing platform to the general public and ongoing collaborative research projects.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Scalable, Massively-Parallel Runtime System with Predictable Performance
Massively Parallel, Adaptive and Fault-Tolerant Execution of Data Flow Programs on Dynamic Clouds
海外基金