Supervised compression of big data

Supervised compression of big data
复制标题

DOI:
10.1002/sam.11508
复制
发表时间:
2021-04
期刊:
Statistical Analysis and Data Mining: The ASA Data Science Journal
影响因子:
--
通讯作者:
V. R. Joseph;Simon Mak
V. R. Joseph;Simon Mak
中科院分区:
其他
文献类型:
--
作者:
V. R. Joseph;Simon Mak

文献摘要

相似文献

大数据现象已经在从科学到工程的几乎所有学科中普遍存在。一个关键的挑战是使用此类数据来拟合统计和机器学习模型,这可能会产生高昂的计算和存储成本。一种解决方案是对精心选择的数据子集执行模型拟合。文献中提出了各种数据缩减方法,从随机二次采样到基于最佳实验设计的方法。然而,当目标是学习潜在的输入输出关系时,这种简化方法可能并不理想,因为它没有利用输出中包含的信息。为此,我们提出了一种称为超级压缩的监督数据压缩方法,该方法通过对对建模所需的输入输出关系最重要的区域进行采样来整合输出信息。超级压缩的优点是它是非参数的——压缩方法不依赖于输入和输出之间的参数建模假设。因此,所提出的方法对于各种建模选择都是稳健的。我们在模拟和出租车预测建模应用中证明了超级压缩相对于现有数据缩减方法的有用性。
The phenomenon of big data has become ubiquitous in nearly all disciplines, from science to engineering. A key challenge is the use of such data for fitting statistical and machine learning models, which can incur high computational and storage costs. One solution is to perform model fitting on a carefully selected subset of the data. Various data reduction methods have been proposed in the literature, ranging from random subsampling to optimal experimental design‐based methods. However, when the goal is to learn the underlying input–output relationship, such reduction methods may not be ideal, since it does not make use of information contained in the output. To this end, we propose a supervised data compression method called supercompress, which integrates output information by sampling data from regions most important for modeling the desired input–output relationship. An advantage of supercompress is that it is nonparametric—the compression method does not rely on parametric modeling assumptions between inputs and output. As a result, the proposed method is robust to a wide range of modeling choices. We demonstrate the usefulness of supercompress over existing data reduction methods, in both simulations and a taxicab predictive modeling application.