Bayesian Dynamic Feature Partitioning in High-Dimensional Regression With Big Data

Bayesian Dynamic Feature Partitioning in High-Dimensional Regression With Big Data
复制标题

大数据高维回归中的贝叶斯动态特征划分

DOI:
10.1080/00401706.2021.1952899
复制
发表时间:
2022
期刊:
影响因子:
2.5
通讯作者:
Guhaniyogi, Rajarshi
Guhaniyogi, Rajarshi
中科院分区:
工程技术3区
文献类型:
--
作者:
Gutierrez, Rene;Guhaniyogi, Rajarshi

文献摘要

参考文献

相似文献

使用马尔可夫链蒙特卡罗 (MCMC) 或其变体对高维线性回归模型进行贝叶斯计算可能非常慢或完全令人望而却步,因为这些方法在采样链的每次迭代中执行昂贵的计算。此外,这种计算成本通常无法在并行架构中有效地划分。如果数据量很大或者数据随着时间的推移按顺序到达(流式传输或在线设置),这些问题会更加严重。本文提出了一种新颖的动态特征分区回归(DFP),用于对大数据或流数据的高维线性回归进行高效在线推理。 DFP 在每个时间点构建参数的伪后验密度,并在新数据块(数据分片)到达时快速更新伪后验密度。 DFP 在每个时间点适当更新伪后验,并对参数集进行分区,以利用并行化实现高效的后验计算。所提出的方法应用于大参数空间上具有高斯尺度混合先验和尖峰和平板先验以及大数据的高维线性回归模型,并且被发现可以产生最先进的推理性能。该算法享有理论支持,随着数据大小的增长,伪后验密度随着时间的推移任意接近完全后验,如补充材料所示。补充材料还包含应用于不同先验的 DFP 算法的详细信息。 https://github.com/Rene-Gutierrez/DynParRegReg 中提供了实施 DFP 广告管理系统的包。该数据集可在 https://github.com/Rene-Gutierrez/DynParRegReg\_Implementation 中获取。
Bayesian computation of high-dimensional linear regression models using Markov chain Monte Carlo (MCMC) or its variants can be extremely slow or completely prohibitive since these methods perform costly computations at each iteration of the sampling chain. Furthermore, this computational cost cannot usually be efficiently divided across a parallel architecture. These problems are aggravated if the data size is large or data arrive sequentially over time (streaming or online settings). This article proposes a novel dynamic feature partitioned regression (DFP) for efficient online inference for high-dimensional linear regressions with large or streaming data. DFP constructs apseudo posterior densityof the parameters at every time point, and quickly updates the pseudo posterior when a new block of data (data shard) arrives. DFP updates the pseudo posterior at every time point suitably and partitions the set of parameters to exploit parallelization for efficient posterior computation. The proposed approach is applied to high-dimensional linear regression models with Gaussian scale mixture priors and spike-and-slab priors on large parameter spaces, along with large data, and is found to yield state-of-the-art inferential performance. The algorithm enjoys theoretical support with pseudoposterior densities over time being arbitrarily close to the full posterior as the data size grows, as shown in the supplementary material. Supplementary material also contains details of the DFP algorithm applied to different priors. Package to implement DFP is available in https://github.com/Rene-Gutierrez/DynParRegReg. The dataset is available in https://github.com/Rene-Gutierrez/DynParRegReg\_Implementation.
DOI: 10.1214/13-aap951
发表时间: 2014-08-01
影响因子: 1.8
作者:
Beskos, Alexandros;Crisan, Dan;Jasra, Ajay
通讯作者: Jasra, Ajay
一般混合物的粒子学习
DOI: 10.1214/10-ba525
发表时间: 2010
期刊: Bayesian Analysis
影响因子: 4.4
作者:
C. Carvalho;H. Lopes;Nicholas G. Polson;Matt Taddy
通讯作者: Matt Taddy
状态空间模型的有偏在线参数推断
DOI: 10.1007/s11009-016-9511-x
发表时间: 2015
影响因子: 0.9
作者:
P. Moral;A. Jasra;Yan Zhou
通讯作者: Yan Zhou
贝叶斯收缩模型的几何遍历性
DOI: 10.1214/14-ejs896
发表时间: 2014
影响因子: 1.1
作者:
Subhadip Pal;K. Khare
通讯作者: K. Khare
DOI: 10.1239/aap/1396360114
发表时间: 2014-03
影响因子: 1.2
作者:
A. Beskos;D. Crisan;A. Jasra;N. Whiteley
通讯作者: A. Beskos;D. Crisan;A. Jasra;N. Whiteley