Big Data Bayesian Linear Regression and Variable Selection by Normal-Inverse-Gamma Summation

Big Data Bayesian Linear Regression and Variable Selection by Normal-Inverse-Gamma Summation
复制标题

DOI:
10.1214/17-ba1083
复制
发表时间:
2017-11
期刊:
影响因子:
4.4
通讯作者:
Hang Qian
Hang Qian
中科院分区:
数学2区
文献类型:
--
作者:
Hang Qian

文献摘要

被引文献

相似文献

。我们引入了正态-逆伽马和算子,它结合了不同数据源的贝叶斯回归结果,得到了一种简单的大数据回归的拆分合并算法。求和算子对于计算边际似然也很有用,并有助于贝叶斯模型选择方法,包括贝叶斯套索、随机搜索变量选择、马尔可夫链蒙特卡罗模型合成等。在一次扫描中,观测值被扫描,然后采样器迭代合并正态-逆伽马分布,而无需重新加载数据。仿真研究表明,我们的算法能够科学地处理高度相关的大数据。ffi。本文还分析了一个关于就业和工资的真实数据集。
. We introduce the normal-inverse-gamma summation operator, which combines Bayesian regression results from different data sources and leads to a simple split-and-merge algorithm for big data regressions. The summation operator is also useful for computing the marginal likelihood and facilitates Bayesian model selection methods, including Bayesian LASSO, stochastic search variable selection, Markov chain Monte Carlo model composition, etc. Observations are scanned in one pass and then the sampler iteratively combines normal-inverse-gamma distributions without reloading the data. Simulation studies demonstrate that our algorithms can efficiently handle highly correlated big data. A real-world data set on employment and wage is also analyzed.