A scalable bootstrap for massive data

A scalable bootstrap for massive data
复制标题

DOI:
10.1111/rssb.12050
复制
发表时间:
2014-09-01
影响因子:
5.8
通讯作者:
Jordan, Michael I.
Jordan, Michael I.
中科院分区:
数学1区
文献类型:
--
作者:
Kleiner, Ariel;Talwalkar, Ameet;Jordan, Michael I.

文献摘要

被引文献

相似文献

自助法提供了一种简单而有力的方法来评估估计量的质量。然而,在涉及大数据集的设置中(这越来越普遍),基于引导的量的计算在计算上可能要求过高。虽然子采样和n中取m自助法等变体原则上可以用于降低自助计算的成本,但这些方法通常对调整参数(例如子采样数据点的数量)的指定不具有鲁棒性,并且与自助法相比,它们通常需要估计量的收敛速度的知识。作为一种替代方案,我们介绍了“小靴带袋”(BLB),这是一个新的程序,它结合了两个引导和二次抽样的功能,以产生一个强大的,计算效率高的方法来评估估计的质量。BLB非常适合于现代并行和分布式计算架构,并且还保留了引导程序的通用适用性和统计效率。我们证明了BLB的有利的统计性能,通过理论分析阐明程序的属性,以及模拟研究比较BLB与引导,m出n引导和二次抽样。此外,我们提出的结果从大规模的分布式实现的BLB展示其计算的优势,在大量的数据,自适应地选择BLB的调整参数的方法,应用BLB到几个真实的数据集和扩展的BLB时间序列数据的实证研究。
The bootstrap provides a simple and powerful means of assessing the quality of estimators. However, in settings involving large data sets-which are increasingly prevalent-the calculation of bootstrap-based quantities can be prohibitively demanding computationally. Although variants such as subsampling and the m out of n bootstrap can be used in principle to reduce the cost of bootstrap computations, these methods are generally not robust to specification of tuning parameters (such as the number of subsampled data points), and they often require knowledge of the estimator's convergence rate, in contrast with the bootstrap. As an alternative, we introduce the 'bag of little bootstraps' (BLB), which is a new procedure which incorporates features of both the bootstrap and subsampling to yield a robust, computationally efficient means of assessing the quality of estimators. The BLB is well suited to modern parallel and distributed computing architectures and furthermore retains the generic applicability and statistical efficiency of the bootstrap. We demonstrate the BLB's favourable statistical performance via a theoretical analysis elucidating the procedure's properties, as well as a simulation study comparing the BLB with the bootstrap, the m out of n bootstrap and subsampling. In addition, we present results from a large-scale distributed implementation of the BLB demonstrating its computational superiority on massive data, a method for adaptively selecting the BLB's tuning parameters, an empirical study applying the BLB to several real data sets and an extension of the BLB to time series data.