Score-matching representative approach for big data analysis with generalized linear models

Score-matching representative approach for big data analysis with generalized linear models
复制标题

DOI:
10.1214/21-ejs1965
复制
发表时间:
2018-07
影响因子:
1.1
通讯作者:
Keren Li;Jie Yang
Keren Li;Jie Yang
中科院分区:
数学3区
文献类型:
--
作者:
Keren Li;Jie Yang

文献摘要

相似文献

我们提出了一种快速有效的策略,称为代表性方法,使用线性模型和广义线性模型进行大数据分析,特别是分布式数据集。该方法在给定大数据集划分的情况下,为每个数据块构造一个代表数据点,并在代表数据集上拟合目标模型。就时间复杂度而言,它与文献中的子采样方法一样快。在效率方面,其参数估计精度优于分治法。此外,当分析分布在不同节点上的海量数据时,代表方法特别有用,因为代表的生成是条件独立的。总的来说,我们推荐两种代表性的方法,均值代表(MR)和分数匹配代表(SMR),沿着理论依据,用于广义线性模型的大数据分析。全面的仿真研究证实,MR是一个很好的解决方案的线性模型和GLM的预分析,而SMR优于子采样和分而治之的方法,即使与中等大小的块,一般GLM。在适当选择数据分区的情况下,SMR估计值似乎甚至与完整数据估计值相当。使用航空公司的准点性能数据作为一个说明性的真实的大数据的例子,我们表明,MR和SMR是一样好的完整的数据估计时,可用的。对于具有平坦的逆链接函数和中等连续瓦里系数的GLM,我们推荐MR。否则,我们推荐MR作为具有更精细划分的初始步骤的SMR解决方案。
We propose a fast and efficient strategy, called the representative approach, with linear models and generalized linear models for big data analysis, and in particular for distributed dataset. With a given partitioning of big dataset, this approach constructs a representative data point for each data block and fits the target model on the representative dataset. In terms of time complexity, it is as fast as the subsampling approaches in the literature. As for effi- ciency, its accuracy of estimated parameters appears to be better than the divide-and-conquer method. Additionally, the representative approach is especially useful when analyzing massive data distributed stored on different nodes, since the generation of representatives is conditional independent. Overall, we recommend two representative approaches, mean representative (MR) and score-matching representative (SMR), along with theoretical justifications, for big data analysis with generalized linear models. Comprehensive simulation studies confirm that MR is a good solution for linear models and pre-analysis for GLMs, while SMR outperforms the subsampling and divide-and-conquer methods, even with moderate size of block, for general GLMs. With properly chosen data partition, SMR estimate appears to be even comparable with the full data estimate. Using the Airline on-time performance data as an illustrative real big data example, we show that MR and SMR are as good as the full data estimate when available. For GLMs with flat inverse link functions and moderate coefficients of the continuous vari- ables, we recommend MR. Otherwise, we recommend SMR solution with MR as an initial step with a finer partition.