Divide-and-Conquer Information-Based Optimal Subdata Selection Algorithm

Divide-and-Conquer Information-Based Optimal Subdata Selection Algorithm
复制标题

DOI:
10.1007/s42519-019-0048-5
复制
发表时间:
2019-05
影响因子:
0.6
通讯作者:
Haiying Wang
Haiying Wang
中科院分区:
--
文献类型:
--
作者:
Haiying Wang

文献摘要

被引文献

相似文献

基于信息的最优子数据选择(IBOSS)是一种 从大数据中选择信息数据点的计算效率高的方法 通过按列处理全部数据来设置。然而,当数据集的体积 太大,无法在机器的可用内存中处理,则不可行 执行IBOSS程序。本文提出了一种分而治之的IBOS 解决这个问题的方法,其中完整的数据集被分成较小的 分区加载到内存中,然后从 每个分区使用IBOSS算法。我们推导出有限样本性质 和所得估计量的渐近性质。渐近结果表明, 如果整个数据集被随机分区而分区的数量不是 非常大,则所得估计器具有与 原始的IBOSS估计器。我们还进行了数值实验,以评估 所提出的方法的经验性能。
The information-based optimal subdata selection (IBOSS) is a computationally efficient method to select informative data points from large data sets through processing full data by columns. However, when the volume of a data set is too large to be processed in the available memory of a machine, it is infeasible to implement the IBOSS procedure. This paper develops a divide-and-conquer IBOSS approach to solving this problem, in which the full data set is divided into smaller partitions to be loaded into the memory and then subsets of data are selected from each partition using the IBOSS algorithm. We derive both finite sample properties and asymptotic properties of the resulting estimator. Asymptotic results show that if the full data set is partitioned randomly and the number of partitions is not very large, then the resultant estimator has the same estimation efficiency as the original IBOSS estimator. We also carry out numerical experiments to evaluate the empirical performance of the proposed method.