Large-sample and deterministic confidence intervals for online aggregation

Large-sample and deterministic confidence intervals for online aggregation
复制标题

DOI:
10.1109/ssdm.1997.621151
复制
发表时间:
1997-08
期刊:
Proceedings. Ninth International Conference on Scientific and Statistical Database Management (Cat. No.97TB100150)
影响因子:
--
通讯作者:
P. Haas
P. Haas
中科院分区:
其他
文献类型:
--
作者:
P. Haas

文献摘要

被引文献

相似文献

J.M. 赫勒斯坦等人(1997年)近期提出的在线聚合系统允许对存储在关系数据库管理系统中的大型复杂数据集进行交互式探索。运行置信区间是在线聚合系统的一个重要组成部分,它向用户表明每个运行聚合与相应最终结果的估计接近程度。大样本置信区间以预先指定的概率包含最终结果,并基于中心极限定理,而确定性置信区间以概率1包含最终查询结果。我们展示了新的和现有的中心极限定理、简单的边界论证以及德尔塔方法如何可用于推导大样本和确定性置信区间的公式。为了说明这些技术,我们在具有连接和选择谓词的单表和多表AVG、COUNT、SUM、VARIANCE和STDEV查询的情况下,获得了运行置信区间的公式。还考虑了去重和GROUP - BY操作。然后,我们提供了用于计算置信区间和分析这些算法复杂性的数值稳定算法。
The online aggregation system recently proposed by J.M. Hellerstein, et al. (1997) permits interactive exploration of large, complex datasets stored in relational database management systems. Running confidence intervals are an important component of an online aggregation system and indicate to the user the estimated proximity of each running aggregate to the corresponding final result. Large sample confidence intervals contain the final result with a prespecified probability and rest on central limit theorems, while deterministic confidence intervals contain the final query result with probability 1. We show how new and existing central limit theorems, simple bounding arguments, and the delta method can be used to derive formulas for both large sample and deterministic confidence intervals. To illustrate these techniques, we obtain formulas for running confidence intervals in the case of single table and multi table AVG, COUNT, SUM, VARIANCE, and STDEV queries with join and selection predicates. Duplicate elimination and GROUP-BY operations are also considered. We then provide numerically stable algorithms for computing the confidence intervals and analyzing the complexity of these algorithms.