An architecture for a data-intensive computer

An architecture for a data-intensive computer
复制标题

数据密集型计算机的体系结构

DOI:
--
复制
发表时间:
2011
期刊:
Network-aware Data Management
影响因子:
--
通讯作者:
R. Burns
R. Burns
中科院分区:
--
文献类型:
--
作者:
E. Givelberg;A. Szalay;Kalin Kanov;R. Burns

文献摘要

被引文献

相似文献

科学仪器和模拟产生了越来越大的数据集,改变了我们进行科学研究的方式。我们提出了一个系统,我们称之为数据密集型计算机,用于对Petascale大小的数据集进行计算。数据密集型计算机由HPC集群、大规模并行数据库和一组运行数据密集型操作系统的计算服务器组成,这将数据库变成数据密集型计算机内存层次结构中的一层。 数据密集型操作系统是面向数据的:传统计算机操作系统的核心是顺序文件的抽象编程模型,取而代之的是对高级数据对象的系统级支持,如多维数组、图形、稀疏数组等。用户应用程序将被编译成在HPC集群和数据库内部执行的代码。然而,数据密集型操作系统是非本地的,允许远程应用程序在数据库内执行代码。该模型支持协作环境,在该环境中,大型数据集通常由大量用户创建和处理。 我们正在开发一个软件库MPI-DB,这是一个数据密集型操作系统的原型。JHU的湍流小组目前正在使用它来将模拟输出存储在数据库中,并执行模拟改进先前存储的结果。
Scientific instruments, as well as simulations, generate increasingly large datasets, changing the way we do science. We propose a system that we call the data-intensive computer for computing with Petascale-sized datasets. The data-intensive computer consists of an HPC cluster, a massively parallel database and a set of computing servers running the data-intensive operating system, which turns the database into a layer in the memory hierarchy of the data-intensive computer. The data-intensive operating system is data-object-oriented: the abstract programming model of a sequential file, central to traditional computer operating systems, is replaced with system-level support for high-level data objects, such as multi-dimensional arrays, graphs, sparse arrays, etc. User application programs will be compiled into code that is executed both on the HPC cluster and inside the database. The data-intensive operating system is however non-local, allowing remote applications to execute code inside the database. This model supports the collaborative environment, where a large data set is typically created and processed by a large group of users. We are developing a software library, MPI-DB, which is a prototype of the data-intensive operating system. It is currently being used by the Turbulence group at JHU to store simulation output in the database and to perform simulations refining previously stored results.