课题基金 / 基金详情

SFB 1404: FONDA - Foundations of Workflows for Large-Scale Scientific Data Analysis

SFB 1404: FONDA - Foundations of Workflows for Large-Scale Scientific Data Analysis
SFB 1404:FONDA - 大规模科学数据分析工作流程的基础
批准号:
414984028
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
德国
项目类别:
Collaborative Research Centres
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
基本上所有的科学学科都在产生越来越多的数据。为了获得科学发现,这些数据集通过复杂的数据分析工作流(Daws)进行分析,这些工作流是一系列离散的分析程序,排列在(通常是非线性)管道中。因为它们通常处理非常大的数据集,所以必须在分布式和/或并行计算基础设施上执行数据仓库。传统上,数据仓库是针对速度进行优化的,这导致解决方案难以复制和共享,并且与一种类型的输入紧密绑定。然而,正如最近的NSF/DOE研讨会总结的那样,该研讨会汇集了工作流程和HPC公报,“......人力生产力可以说仍然是最昂贵的资源,胜过电力,性能,拟议的CRC方达-“大规模科学数据分析工作流程基础”-将继续进行这一观察,并研究提高大型科学数据集的数据仓库开发、执行和维护效率的方法。我们的长期目标是开发方法和工具,实现大大减少开发时间和开发成本的数据仓库。我们将从一个基本的角度来处理这些问题,即,我们的目标是找到新的抽象、模型和算法,这些抽象、模型和算法最终可以形成新的一类未来可编程基础设施的基础。为了实现这些目标,方达在其第一阶段将着重于数据采集和分析引擎的三个关键特性,即便携性、适应性和可靠性。我们希望研究以下问题的答案:我们如何构建能够在不同基础设施之间实现分析可移植性的DAO和数据挖掘引擎?必须如何设计数据仓库以适应不断变化的输入数据或略有变化的需求?我们如何才能建立可靠的可持续发展系统,意识到并控制自己的限制和前提条件?数据仓库是两个世界之间的桥梁:第一,使用计算机的特定科学学科,第二,计算机科学,它建立了开发和执行数据仓库所需的基础设施。因此,发展新的科学数据仓库基础需要这两个世界之间的密切互动。方达通过建立计算机科学、材料科学、地球科学和生命科学的跨学科PI小组来实现这一想法。通过这些合作,方达的研究成果将不断得到验证,利用自然科学不同领域的相关和当前的科学问题。
英文摘要
Essentially all scientific disciplines are generating an ever-increasing amount of data. To derive scientific discoveries, these data sets are analyzed by complex data analysis workflows (DAWs), which are series of discrete analysis programs arranged in (often non-linear) pipelines. Because they usually deal with very large data sets, DAWs must be executed on distributed and/or parallel computational infrastructures. Traditionally, DAWs are optimized for speed, which leads to solutions that are hard to reproduce and share and that are tightly bound to exactly one type of input. However, as stated as summary in a recent NSF/DOE workshop that brought together the workflow and the HPC communi-ties, “… human productivity arguably still is the most expensive resource, trumping power, perfor-mance, and other factors …”.The proposed CRC FONDA – “Foundations of workflows for large-scale scientific data analysis” – will take up this observation and investigate methods for increasing productivity in the development, execution, and maintenance of DAWs for large scientific data sets. Our long-term goal is to develop methods and tools that achieve substantial reductions in development time and development cost of DAWs. We will approach these questions from a fundamental perspective, i.e., we aim at finding new abstractions, models, and algorithms that can eventually form the basis of a new class of future DAW infrastructures. Toward these goals, FONDA in its first phase will focus on three critical properties of DAWs and of DAW engines, namely portability, adaptability, and dependability (PAD). We want to investigate answers to questions such as: How can we build DAWs and DAW engines that enable portability of analysis across different infrastructures? How must DAWs be designed to adapt to changing input data or slightly changing requirements? How can we build dependable DAW systems that are aware of and control their own limitations and preconditions?DAWs are bridges between two worlds: First, the specific scientific discipline using a DAW, and, sec-ond, Computer Science, which builds the infrastructures necessary for developing and executing DAWs. Developing novel foundations for scientific DAWs thus requires a close interaction between these two worlds. FONDA implements this idea by building on an interdisciplinary group of PIs from Computer Science, Material Science, Geosciences, and the Life Sciences. Through these cooperations, FONDA’s research results will be continuously validated using relevant and current scientific problems from different fields of the natural sciences.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金