Simulating Multi-Tenant OLAP Database Clusters

Simulating Multi-Tenant OLAP Database Clusters
复制标题

DOI:
--
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
J. Schaffner;Benjamin Eckart;Christian Schwarz;Jan Brunnert;D. Jacobs;A. Zeier;H. Plattner
J. Schaffner;Benjamin Eckart;Christian Schwarz;Jan Brunnert;D. Jacobs;A. Zeier;H. Plattner
中科院分区:
其他
文献类型:
--
作者:
J. Schaffner;Benjamin Eckart;Christian Schwarz;Jan Brunnert;D. Jacobs;A. Zeier;H. Plattner

文献摘要

被引文献

相似文献

在20世纪90年代,并行数据库机的模拟被用于许多数据库研究项目。模拟方法在当时流行的主要原因之一是,具有数百个节点的集群并不像今天这样容易进行实验。同时,这些系统的仿真模型相当复杂,因为它们需要在硬件(例如CPU争用或磁盘I/O)和软件(例如处理分布式连接)中捕获排队过程。如今,数据库体系结构趋向于更加专业化,这就消除了建模任务中的大部分复杂性。作为本文的主要贡献,我们讨论了如何开发一个简单的模拟模型,这样一个专门的系统:一个多租户的OLAP集群的基础上,在内存中的列数据库。最初的基础设施和测试平台是使用SAP TREX构建的,这是SAP业务仓库加速器的内存列数据库部分,我们将其移植到Amazon EC2云上运行。虽然我们采用了一个简单的排队模型,我们取得了良好的准确性。类似于20世纪90年代的一些并行系统,我们有兴趣在仿真的帮助下研究不同的复制和高可用性策略。特别是,我们研究了镜像与交错复制的吞吐量和负载分布在我们的集群的多租户数据库的影响。我们表明,更好的负载分布固有的交错复制策略表现在EC2和我们的模拟环境。
Simulation of parallel database machines was used in many database research projects during the 1990ies. One of the main reasons why simulation approaches were popular in that time was the fact that clusters with hundreds of nodes were not as readily available for experimentation as it is the case today. At the same time, the simulation models underlying these systems were fairly complex since they needed to capture both queuing processes in hardware (e.g. CPU contention or disk I/O) and software (e.g. processing distributed joins). Todays trend towards more specialized database architectures removes large parts of this complexity from the modeling task. As the main contribution of this paper, we discuss how we developed a simple simulation model of such a specialized system: a multi-tenant OLAP cluster based on an in-memory column database. The original infrastructure and testbed was built using SAP TREX, an in-memory column database part of SAP’s business warehouse accelerator, which we ported to run on the Amazon EC2 cloud. Although we employ a simple queuing model, we achieve good accuracy. Similar to some of the parallel systems of the 1990ies, we are interested in studying different replication and high-availability strategies with the help of simulation. In particular, we study the effects of mirrored vs. interleaved replication on throughput and load distribution in our cluster of multi-tenant databases. We show that the better load distribution inherent to the interleaved replication strategy is exhibited both on EC2 and in our simulation environment.