The partitioned exponential file for database storage management

The partitioned exponential file for database storage management
复制标题

DOI:
10.1007/s00778-005-0171-7
复制
发表时间:
2007-10
期刊:
The VLDB Journal
影响因子:
--
通讯作者:
C. Jermaine;E. Omiecinski;Wai Gen Yee
C. Jermaine;E. Omiecinski;Wai Gen Yee
中科院分区:
其他
文献类型:
--
作者:
C. Jermaine;E. Omiecinski;Wai Gen Yee

文献摘要

被引文献

相似文献

硬盘存储容量的增长率继续超过硬盘寻道时间的下降率。这种趋势意味着寻道的值相对于存储的值呈指数级增长。考虑到这种趋势,我们引入了分区指数文件(PE文件),这是一种通用的存储管理器,可以为许多不同类型的数据(例如,数值的、空间的或时间的)。PE文件适用于具有密集更新负载和并发分析查询的环境。例如,在可以产生PB级数据的长期运行的科学应用程序中可以找到这样的环境。例如,拟议中的大型综合巡天望远镜[36]将在其多年的寿命中产生50-100 PB的观测和科学数据。这个数据库永远不会离线,因此每天数十TB的突发更新负载必须与数据分析同时处理。在PE文件中,数据被组织为一系列磁盘上的排序,并进行了仔细的全局组织。由于PE文件在很大程度上依赖于顺序I/O,只有一小部分的磁盘寻求是需要一个典型的记录插入或retrieval. FIGURE描述的PE文件,我们还详细介绍了一组基准测试实验forT 1 SM,这是一个PE文件定制使用的多属性数据记录在一个单一的数值属性。在我们的基准测试中,我们实现并测试了许多可用于索引和存储此类数据的竞争数据组织,例如B+树、LSM-Tree、Buffer Tree、Stepped Merge Method和Y-Tree。正如预期的那样,没有一个组织是所有基准测试中最好的,但我们的实验表明,T1 SM在许多情况下是最好的选择,这表明它是最好的整体。具体来说,T1 SM在必须与密集插入流并发处理的重查询工作负载的情况下表现得非常好。我们的实验表明,T1 SM(及其近亲,T2 SM空间数据存储管理器)可以处理这种类型的非常繁重的混合工作负载,并仍然保持可接受的小查询延迟。
The rate of increase in hard disk storage capacity continues to outpace the rate of decrease in hard disk seek time. This trend implies that the value of a seek is increasing exponentially relative to the value of storage.With this trend in mind, we introduce thepartitioned exponential file(PE file) which is a generic storage manager that can be customized for many different types of data (e.g., numerical, spatial, or temporal). The PE file is intended for use in environments with intense update loads and concurrent, analytic queries. Such an environment may be found, for example, in long-running scientific applications which can produce petabytes of data. For example, the proposed Large Synoptic Survey Telescope [36] will produce 50–100 petabytes of observational, scientific data over its multi-year lifetime. This database will never be taken off-line, so bursty update loads of tens of terabytes per day must be handled concurrently with data analysis. In the PE file, data are organized as a series of on-disk sorts with a careful, global organization. Because the PE file relies heavily on sequential I/O, only a fraction of a disk seek is required for a typical record insertion or retrieval.In addition to describing the PE file, we also detail a set of benchmarking experiments forT1SM, which is a PE file customized for use with multi-attribute data records ordered on a single numerical attribute. In our benchmarking, we implement and test many competing data organizations that can be used to index and store such data, such as the B+-Tree, the LSM-Tree, the Buffer Tree, the Stepped Merge Method, and the Y-Tree. As expected, no organization is the best over all benchmarks, but our experiments show that T1SM is the best choice in many situations, suggesting that it is the best overall. Specifically, T1SM performs exceptionally well in the case of a heavy query workload that must be handled concurrently with an intense insertion stream. Our experiments show that T1SM (and its close cousin, the T2SM storage manager for spatial data) can handle very heavy mixed workloads of this type, and still maintain acceptably small query latencies.