On-the-Fly Data Synopses Efficient Data Exploration in the Simulation Sciences

On-the-Fly Data Synopses Efficient Data Exploration in the Simulation Sciences
复制标题

动态数据概要模拟科学中的高效数据探索

DOI:
10.1145/2814710.2814715
复制
发表时间:
2015
期刊:
ACM SIGMOD Record
影响因子:
--
通讯作者:
Heinis T
Heinis T
中科院分区:
--
文献类型:
--
作者:
Heinis T

文献摘要

相似文献

由于计算硬件越来越强大,仪器越来越精密,我们产生科学数据的能力远远超过了有效存储和分析数据的能力。当今的科学数据分析工具很少能够处理由仪器捕获或由超级计算机生成的海量数据。然而,在许多情况下,详细分析一小部分数据就足够了。因此,分析数据的科学家需要的是使用近似查询结果来探索完整数据集并识别感兴趣的子集的有效方法。一旦发现了有趣的区域,仍然可以使用精确但更耗时的分析来仔细检查。数据概要符合要求,因为它们提供了对大量数据的快速(但近似)查询执行。然而,在数据存储之后生成数据概要要求我们再次分析所有数据,因此效率很低。我们建议在捕获数据时实时生成模拟应用程序的概要。这样做通常意味着更改模拟或数据捕获代码,这是乏味的,通常只是一次性的解决方案,并不普遍适用。相比之下,我们的愿景为科学家提供了一种高级语言和基础设施,以便在模拟运行时生成实时创建数据概要的代码。在本文中,我们将讨论与我们的方法相关的数据管理挑战
As a consequence of ever more powerful computing hardware and increasingly precise instruments, our capacity to produce scientific data by far outpaces our ability to efficiently store and analyse it. Few of today's tools to analyse scientific data are able to handle the deluge captured by instruments or generated by supercomputers.In many scenarios, however, it suffices to analyse a small subset of the data in detail. What scientists analysing the data consequently need are efficient means to explore the full dataset using approximate query results and to identify the subsets of interest. Once found, interesting areas can still be scrutinised using a precise, but also more time-consuming analysis. Data synopses fit the bill as they provide fast (but approximate) query execution on massive amounts of data. Generating data synopses after the data is stored, however, requires us to analyse all the data again, and is thus inefficientWhat we propose is to generate the synopsis for simulation applications on-the-fly when the data is captured. Doing so typically means changing the simulation or data capturing code and is tedious and typically just a one-off solution that is not generally applicable. In contrast, our vision gives scientists a high-level language and the infrastructure needed to generate code that creates data synopses on-the-fly, as the simulation runs. In this paper we discuss the data management challenges associated with our approach