Experimental evidence on partitioning in parallel data warehouses

Experimental evidence on partitioning in parallel data warehouses
复制标题

并行数据仓库分区的实验证据

DOI:
10.1145/1031763.1031769
复制
发表时间:
2004
期刊:
2016 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
P. Furtado
P. Furtado
中科院分区:
--
文献类型:
--
作者:
P. Furtado

文献摘要

被引文献

相似文献

并行性可用于具有性能和可伸缩性挑战的大型数据仓库(DW)的重大性能改进。一个简单的低成本共享架构具有水平分区的事实,可用于大大加速数据仓库的响应时间。但是,如果在放置期间不采取特殊护理以最大程度地减少此类开销,则与节点之间的大型复制关系和重新分配要求有关的额外开销会大大降低加速性能。在本文中,我们在绩效评估基准TPC-H的帮助下在实验中显示了这些问题,并确定了可以最大程度地减少这种不受欢迎的额外开销的简单修改。我们在实验上分析了一个简单易用的分区和放置决策,可实现良好的性能改进结果。
Parallelism can be used for major performance improvement in large Data warehouses (DW) with performance and scalability challenges. A simple low-cost shared-nothing architecture with horizontally fully-partitioned facts can be used to speedup response time of the data warehouse significantly. However, extra overheads related to processing large replicated relations and repartitioning requirements between nodes can significantly degrade speedup performance for many query patterns if special care is not taken during placement to minimize such overheads. In this paper we show these problems experimentally with the help of the performance evaluation benchmark TPC-H and identify simple modifications that can minimize such undesirable extra overheads. We analyze experimentally a simple and easy-to-apply partitioning and placement decision that achieves good performance improvement results.