An approach for automatic data virtualization

An approach for automatic data virtualization
复制标题

一种自动数据虚拟化方法

DOI:
--
复制
发表时间:
2004
期刊:
Proceedings. 13th IEEE International Symposium on High performance Distributed Computing, 2004.
影响因子:
--
通讯作者:
J. Saltz
J. Saltz
中科院分区:
--
文献类型:
--
作者:
L. Weng;G. Agrawal;Ümit V. Çatalyürek;T. Kurç;S. Narayanan;J. Saltz

文献摘要

被引文献

相似文献

大型和/或地理上分布的科学数据集的分析正在成为网格计算的关键组成部分。这一领域的一个挑战是科学数据集通常存储为二进制或字符平面文件,这使得处理规范更加困难。鉴于此,近来对数据虚拟化和支持这种虚拟化的数据服务产生了兴趣。本文提出了一种自动创建数据服务以支持数据虚拟化的方法。具体来说,我们展示了如何像数据抽象的关系表可以支持复杂的多维科学数据集,驻留在一个集群。我们已经设计并实现了一个工具,可以在多维数据集上处理SQL查询(使用select和where语句)。我们设计了一个元数据描述语言,用于指定数据布局。从这样的描述,我们的工具自动生成高效的数据子集和访问功能。我们已经对我们的系统进行了广泛的评估。我们实验的主要观察结果如下。首先,我们的工具可以正确有效地处理各种不同的数据布局。其次,我们的系统可以随着节点数量或数据量的扩展而扩展。第三,自动生成的索引和收缩功能代码的性能与手写代码的性能相当。
Analysis of large and/or geographically distributed scientific datasets is emerging as a key component of grid computing. One challenge in this area is that scientific datasets are typically stored as binary or character flat-files, which makes specification of processing much harder. In view of this, there has been recent interest in data virtualization, and data services to support such virtualization. This paper presents an approach for automatically creating data services to support data virtualization. Specifically, we show how a relational table like data abstraction can be supported for complex multidimensional scientific datasets that are resident on a cluster. We have designed and implemented a tool that processes SQL queries (with select and where statements) on multi-dimensional datasets. We have designed a meta-data description language that is used for specifying the data layout. From such description, our tool automatically generates efficient data subsetting and access functions. We have extensively evaluated our system. The key observations from our experiments are as follows. First, our tool can correctly and efficiently handle a variety of different data layouts. Second, our system scales well as the number of nodes or the amount of data is scaled. Third, the performance of the automatically generated code for indexing and contracting functions is quite comparable to the performance of hand-written codes.