Sequential sampling procedures for query size estimation

Sequential sampling procedures for query size estimation
复制标题

用于估计查询大小的顺序采样过程

DOI:
--
复制
发表时间:
1992
期刊:
ACM SIGMOD Conference
影响因子:
--
通讯作者:
A. Swami
A. Swami
中科院分区:
--
文献类型:
--
作者:
P. Haas;A. Swami

文献摘要

被引文献

相似文献

我们提供了一个基于随机抽样的过程,用于估计查询结果的大小。该过程是顺序的,因为根据取决于到目前为止获得的观测值的停止规则,采样在随机数目的步骤之后终止。获得足够的观测值,使得在具有预定概率的情况下,估计值与查询结果的真实大小相差不超过预定的量。与以前的查询序贯估计过程不同,我们的过程是渐近有效的,并且不需要特别的试点样本或关于数据特征的先验假设。除了建立估计过程的渐近性质外,我们还提供了在小样本量下减少欠覆盖的技术,并证明了通过分层抽样技术可以降低估计过程的抽样成本。
We provide a procedure, based on random sampling, for estimation of the size of a query result. The procedure is sequential in that sampling terminates after a random number of steps according to a stopping rule that depends upon the observations obtained so far. Enough observations are obtained so that, with a pre-specified probability, the estimate differs from the true size of the query result by no more than a prespecified amount. Unlike previous sequential estimation procedures for queries, our procedure is asymptotically efficient and requires no ad hoc pilot sample or a a priori assumptions about data characteristics. In addition to establishing the asymptotic properties of the estimation procedure, we provide techniques for reducing undercoverage at small sample sizes and show that the sampling cost of the procedure can be reduced through stratified sampling techniques.