A supplement to sampling-based methods for query size estimation in a database system

A supplement to sampling-based methods for query size estimation in a database system
复制标题

对数据库系统中基于采样的查询大小估计方法的补充

DOI:
--
复制
发表时间:
1992
期刊:
SGMD
影响因子:
--
通讯作者:
Wei Sun
Wei Sun
中科院分区:
--
文献类型:
--
作者:
Y. Ling;Wei Sun

文献摘要

被引文献

相似文献

近年来,基于抽样的关系算子(如选择、连接和投影)后估计关系大小的方法得到了广泛的研究。这类方法可以达到较高的估计精度和效率。由于基于抽样的方法涉及的主要开销是抽样成本,因此提出了不同的抽样方法变种,以最小化抽样百分比(从而降低抽样成本),同时保持在置信度和相对误差方面的估计精度(将在第2节后面精确定义)。为了确定最小抽样百分比,需要了解数据的总体特征,如均值和方差。目前,文献中代表性的基于抽样的方法都是基于数据的总体特征不可用的假设,因此大量的工作致力于估计这些特征以接近最优(最小)抽样百分比。对这些特征的估计不仅会产生估计误差,而且会产生成本。在这篇短文中,我们指出,这些数据特征的精确值可以在数据库系统中跟踪,开销可以忽略不计。结果,可以精确地确定在确保指定的相对误差和置信度的情况下的最小抽样百分比。
Sampling-based methods for estimating relation sizes after relational operators such as selections, joins and projections have been intensively studied in recent years. Methods of this type can achieve high estimation accuracy and efficiency. Since the dominating overhead involved in a sampling-based method is the sampling cost, different variants of sampling methods are proposed so as to minimize the sampling percentage (thus reducing the sampling cost) while maintaining the estimation accuracy in terms of the confidence level and relative error (to be precisely defined later in Section 2). In order to determine the minimal sampling percentage, the overall characteristics of the data such as the mean and variance are needed. Currently, the representative sampling-based methods in literature are based on the assumption that overall characteristics of data are unavailable, and thus a significant amount of effort is dedicated to estimating these characteristics so as to approach the optimal (minimal) sampling percentage. The estimation for these characteristics incurs cost as well as suffers the estimation error. In this short essay, we point out that the exact values of these characteristics of data can be kept track of in a database system at a negligible overhead. As a result, the minimal sampling percentage while ensuring the specified relative error and confidence level can be precisely determined.