Analytics-driven Efficient Indexing and Query Processing of Extreme Scale AMR Data
Analytics-driven Efficient Indexing and Query Processing of Extreme Scale AMR Data
批准号:
1240682
负责人:
Nagiza Samatova
金额:
$15.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-05-01 至 2015-12-31
中文摘要
对自然现象的高保真科学模拟,如大规模恒星爆炸、飓风运动或聚变能源生产,越来越不仅需要大量的计算,而且需要大量的数据。大规模模拟产生的“大数据”需要可持续的数据管理解决方案,以提高能源效率,实现预测性数据分析,并减少发现时间。为响应这一迫切需求,自适应网格细化(AMR)作为一种很有前途的技术应运而生,它在保持甚至提高模拟精度的同时,实现了相当大的内存、计算和存储资源的节省。有了AMR,科学家可以自适应地在高分辨率观察模拟现象的“不寻常”区域和低精度扫视模拟现象的“通常”区域之间取得平衡。AMR对模拟的显著区域进行选择性细粒度细化的优势也成为其弱点。其底层网格的非统一结构导致在从数据分析到科学发现的过程中对“大数据”的模拟后访问效率低下。时空模拟的数据生成过程从一个时间步前进到下一个时间步,并且只需要两个时间步的上下文,同时只在盘上存储一个时间步的数据。相比之下,可视化数据分析通常需要可用数据的完整上下文,而不仅仅是单个时间步长。事实上,由局部时空关系驱动的模拟在很大程度上是为了通过交互式的假设数据探索来发现或解释非局部和大规模的时空关系。因此,数据环境的根本差异和访问模式的异构性呼唤分析驱动的科学数据管理解决方案。为了弥合这一差距,该项目旨在通过使数据分析和数据缩减成为数据库设计和查询处理的一等公民,显著推进对“大”AMR数据的数据库管理。为了支持分析驱动的高效查询处理,它提供了从传统的数据索引到建议的关于数据压缩和存储中的分层数据布局的信息索引的转变。前者产生大量存储开销,而后者预计存储友好,响应用户查询的数据检索速度显著加快。为了实现这一雄心勃勃的目标,在数据密集型计算机科学的三个广泛领域提出了创新:语义建模、科学数据库管理和高性能存储。具体地说,科学数据库管理将认识到特定于AMR数据的丰富的数据访问和查询语义,能够通过模拟在数据生成过程中对数据进行现场压缩和索引,并通过新的数据布局方法针对不同的访问模式进行优化。这项研究是与三个国家实验室的研究团队密切合作进行的,这些团队在高性能存储、AMR和高保真气候变化模拟方面拥有领先技术。这项研究中开发的方法和形式预计将适用于广泛的科学和工程问题,这些问题利用模拟模型通过“大数据”分析对物理过程进行预测性理解。在该项目下开发的技术预计将加快科学工作流程中以数据为导向的探索和知识发现的关键进程。该项目还计划在“大数据”教育、多样性、社区参与以及通过学术出版物和开源软件传播成果方面做出重大努力。
英文摘要
High-fidelity scientific simulations of natural phenomena, such as massive star explosion, hurricane movement, or fusion energy production increasingly become not only compute-intensive but also data-intensive. "Big" data produced by large-scale simulations calls for sustainable data management solutions that increase energy efficiency, enable predictive data analytics, and reduce time-to-discovery. In response to this critical need, Adaptive Mesh Refinement (AMR) has emerged as a promising technology to achieve considerable savings in memory, computation, and storage resources, while maintaining or even increasing simulation accuracy. With AMR, scientists can adaptively balance between high-resolution look into "the unusual" areas and low-precision glance at "the usual" areas of the simulated phenomena.The strength of AMR selective fine-grain refinement of salient areas of the simulationis also becoming its weakness. The non-uniform structure of its underlying mesh leads to inefficient post-simulation access to the "Big" data along the path from data analytics to scientific discovery. The data generation process of space-time simulation proceeds from one time step to the next and requires the context of only two time steps, while storing data for only one time step on the disk. In contrast, visual data analytics often requires the full context of the available data, not just a single time step. In fact, simulations that are driven by local space-time relationships are largely performed with the purpose of discovering or explaining non-local and large-scale space-time relationships through interactive "what-if" data exploration. Thus, the fundamental differences in data context and heterogeneity of access patterns call for analytics-driven scientific data management solutions.To close this gap, this project aims to significantly advance database management for "Big" AMR data by making data analytics and data reduction the first class citizens of the database design and query processing. To support analytics-driven efficient query processing, it offers a transformative shift from the traditional indexing of data to the proposed indexing of information about data compression and hierarchical data layout in storage. The former incurs substantial storage overhead, while the latter is anticipated to be storage-friendly with significantly faster data retrieval in response to user queries. To realize this ambitious goal, innovations are proposed in three broad areas of data-intensive computer science: semantic modeling, scientific database management, and high performance storage. Specifically, the scientific database management is going to be cognizant of rich data access-and-query semantics specific to AMR data, capable of in situ data compression and indexing during data generation by the simulation, and optimized for heterogeneous access patterns through novel data layout methodologies. This research is being conducted in a close partnership with the three national labs' research teams with pioneering technologies in high performance storage, AMR, and high-fidelity climate change simulations. The approaches and formalisms developed in this research are expected to be applicable to a broad range of scientific and engineering problems that utilize simulation models for predictive understanding of physical processes through "Big" data analytics. The techniques developed under this project are expected to accelerate the crucial process of data-driven exploration and knowledge discovery in the scientific workflow. This project also plans significant efforts in "Big" data education, diversity, community engagement, and the dissemination of results through academic publications and open-source software.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Understanding Climate Change: A Data Driven Approach
-
批准号:1028746
-
项目类别:Continuing Grant
-
资助金额:$179.97万
-
财政年份:2010
-
负责人:Nagiza Samatova
-
依托单位:
Workshop on Mathematics for Petascale Data, June 3-5, 2008, Rockville, MD
-
批准号:0829830
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2008
-
负责人:Nagiza Samatova
-
依托单位:
国内基金
海外基金
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
基于Cache的远程计时攻击研究
-
批准号:60772082
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2007
-
负责人:王韬
-
依托单位: