DCMS: A data analytics and management system for molecular simulation.

DCMS: A data analytics and management system for molecular simulation.
复制标题

DOI:
10.1186/s40537-014-0009-5
复制
发表时间:
2015
影响因子:
8.1
通讯作者:
Xia Y
Xia Y
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kumar A;Grupcev V;Berrada M;Fogarty JC;Tu YC;Zhu X;Pandit SA;Xia Y

文献摘要

相似文献

分子模拟是研究大系统物理化学特性的有力工具,在许多科学和工程领域得到了广泛的应用。在模拟过程中,实验产生了大量的原子,并打算观察它们的空间和时间关系,以进行科学分析。庞大的数据量及其密集的交互给数据访问、管理和分析带来了重大挑战。到目前为止,现有的MS软件系统在存储和处理MS数据方面存在不足,主要是因为缺少一个平台来支持涉及密集数据访问和分析过程的应用程序。在本文中,我们介绍了我们的团队在过去几年中开发的数据库为中心的分子模拟(DCMS)系统。DCMS背后的主要思想是将MS数据存储在关系数据库管理系统(DBMS)中,以利用声明性查询接口(即,SQL)、数据访问方法、查询处理和现代DBMS的优化机制。一个独特的挑战是处理通常是计算密集型的分析查询。为此,我们开发了新的索引和查询处理策略(包括在现代协处理器上运行的算法)作为DBMS的集成组件。因此,研究人员可以使用DBMS内部实现的高效功能上传和分析数据。生成索引结构以存储其他用户可能感兴趣的分析结果,使得结果在不重复分析的情况下容易获得。我们已经开发了一个原型的DCMS的PostgreSQL系统的基础上,使用真实的MS数据和工作负载的实验表明,DCMS显着优于现有的MS软件系统。我们还将其用作测试其他数据管理问题(如安全性和压缩)的平台。
Molecular Simulation (MS) is a powerful tool for studying physical/chemical features of large systems and has seen applications in many scientific and engineering domains. During the simulation process, the experiments generate a very large number of atoms and intend to observe their spatial and temporal relationships for scientific analysis. The sheer data volumes and their intensive interactions impose significant challenges for data accessing, managing, and analysis. To date, existing MS software systems fall short on storage and handling of MS data, mainly because of the missing of a platform to support applications that involve intensive data access and analytical process. In this paper, we present the database-centric molecular simulation (DCMS) system our team developed in the past few years. The main idea behind DCMS is to store MS data in a relational database management system (DBMS) to take advantage of the declarative query interface (i.e., SQL), data access methods, query processing, and optimization mechanisms of modern DBMSs. A unique challenge is to handle the analytical queries that are often compute-intensive. For that, we developed novel indexing and query processing strategies (including algorithms running on modern co-processors) as integrated components of the DBMS. As a result, researchers can upload and analyze their data using efficient functions implemented inside the DBMS. Index structures are generated to store analysis results that may be interesting to other users, so that the results are readily available without duplicating the analysis. We have developed a prototype of DCMS based on the PostgreSQL system and experiments using real MS data and workload show that DCMS significantly outperforms existing MS software systems. We also used it as a platform to test other data management issues such as security and compression.