Estimating aggregates in time-constrained approximate queries in Oracle

Estimating aggregates in time-constrained approximate queries in Oracle
复制标题

估计 Oracle 中时间受限的近似查询中的聚合

DOI:
--
复制
发表时间:
2009
期刊:
International Conference on Extending Database Technology
影响因子:
--
通讯作者:
Jagannathan Srinivasan
Jagannathan Srinivasan
中科院分区:
--
文献类型:
--
作者:
Ying Hu;S. Sundara;Jagannathan Srinivasan

文献摘要

被引文献

相似文献

为了解决长时间运行的SQL查询问题,引入了时间受限的SQL查询的概念。支持时间受限的SQL查询的一个关键方法是使用采样来减少需要处理的数据量,从而允许在指定的时间约束内完成查询。然而,采样确实使查询结果接近,因此需要系统估计出现在选择列表中的表达式(尤其是聚合)的值。因此,对于时间受限的近似SQL查询来说,提出聚合的估计是至关重要的,这也是本文的重点。具体地说,我们解决了在时间受限的近似查询中估计经常出现的聚集(即SUM、COUNT、AVG、MENTAL、MIN和MAX)的问题。对于不同类型的查询,我们使用Bernoulli抽样给出了关于SUM、COUNT、AVG和中值的点估计和区间估计,包括使用叉积抽样的连接处理。对于MIN(MAX),我们给出了100γ%的总体比例将超过从采样数据获得的MIN(或小于MAX)的置信度。
The concept of time-constrained SQL queries was introduced to address the problem of long-running SQL queries. A key approach adopted for supporting time-constrained SQL queries is to use sampling to reduce the amount of data that needs to be processed, thereby allowing completion of the query in the specified time constraint. However, sampling does make the query results approximate and hence requires the system to estimate the values of the expressions (especially aggregates) occurring in the select list. Thus, coming up with estimates for aggregates is crucial for time-constrained approximate SQL queries to be useful, which is the focus of this paper. Specifically, we address the problem of estimating commonly occurring aggregates (namely, SUM, COUNT, AVG, MEDIAN, MIN, and MAX) in time-constrained approximate queries. We give both point and interval estimates for SUM, COUNT, AVG, and MEDIAN using Bernoulli sampling for various type of queries, including join processing with cross product sampling. For MIN (MAX), we give the confidence level that the proportion 100γ% of the population will exceed the MIN (or be less than the MAX) obtained from the sampled data.