Faster methods for random sampling

Faster methods for random sampling
复制标题

DOI:
10.1145/358105.893
复制
发表时间:
1984-07
期刊:
Commun. ACM
影响因子:
--
通讯作者:
J. Vitter
J. Vitter
中科院分区:
其他
文献类型:
--
作者:
J. Vitter

文献摘要

被引文献

相似文献

提出了几种新方法,以随机选择n记录,而无需从包含n记录的文件中替换。在没有预处理的情况下选择样本的记录。 d会在o(n)时间中进行采样,在抽样过程中,大致生成了n均匀的随机变化,大约是n个凸起操作(对实数A和B形式)在大型计算机上解决了文献中的一个开放问题。
Several new methods are presented for selecting n records at random without replacement from a file containing N records. Each algorithm selects the records for the sample in a sequential manner—in the same order the records appear in the file. The algorithms are online in that the records for the sample are selected iteratively with no preprocessing. The algorithms require a constant amount of space and are short and easy to implement. The main result of this paper is the design and analysis of Algorithm D, which does the sampling in O(n) time, on the average; roughly n uniform random variates are generated, and approximately n exponentiation operations (of the form ab, for real numbers a and b) are performed during the sampling. This solves an open problem in the literature. CPU timings on a large mainframe computer indicate that Algorithm D is significantly faster than the sampling algorithms in use today.