An Approach to Incorporate Subsampling into a Generic Bayesian Hierarchical Model.

An Approach to Incorporate Subsampling into a Generic Bayesian Hierarchical Model.
复制标题

DOI:
10.1080/10618600.2021.1923518
复制
发表时间:
2021
期刊:
Journal of computational and graphical statistics : a joint publication of American Statistical Association, Institute of Mathematical Statistics, Interface Foundation of North America
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

本文的目标是提供一种方法,贝叶斯统计学家将子抽样直接到贝叶斯层次模型的选择,而不施加额外的限制性模型假设。我们的动机是,“大数据”的兴起给统计学家直接将其方法应用于大数据集带来了困难。我们引入了一个“数据子集模型”流行的“数据模型,过程模型,参数模型”的框架,用于总结贝叶斯分层模型。数据子集模型的超参数被构造性地指定,因为它们被选择为使得子集的隐含大小满足预定义的计算约束。因此,这些超参数有效地将统计模型校准到计算机本身,以在预先指定的时间量内获得预测/估计。提供了数据子集模型的几个属性,包括:适当性,部分充分性和半参数属性。模拟数据集将用于评估二次抽样的结果,结果将在不同的计算机上显示,以显示计算机对统计分析的影响。此外,我们还提供了一个高维数据集(大约10千兆字节)的联合分析,该数据集由美国人口普查局公共使用微样本(Public Use Micro-Sample,简称CAMS)的2018年5年期估计数据组成。
The goal of this paper is to provide a way for Bayesian statisticians to incorporate subsampling directly into the Bayesian hierarchical model of their choosing without imposing additional restrictive model assumptions. We are motivated by the fact that the rise of “big data” has created difficulties for statisticians to directly apply their methods to big datasets. We introduce a “data subset model” to the popular “data model, process model, and parameter model” framework used to summarize Bayesian hierarchical models. The hyperparameters of the data subset model are specified constructively in that they are chosen such that the implied size of the subset satisfies pre-defined computational constraints. Thus, these hyperparameters effectively calibrate the statistical model to the computer itself to obtain predictions/estimations in a pre-specified amount of time. Several properties of the data subset model are provided including: propriety, partial sufficiency, and semi-parametric properties. Simulated datasets will be used to assess the consequences of subsampling, and results will be presented across different computers to show the effect of the computer on the statistical analysis. Additionally, we provide a joint analysis of a high-dimensional dataset (roughly 10 gigabytes) consisting of 2018 5-year period estimates from the US Census Bureau’s Public Use Micro-Sample (PUMS).
DOI: 10.1093/biomet/asr054
发表时间: 2011-12-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
Bien, Jacob;Tibshirani, Robert J.
通讯作者: Tibshirani, Robert J.
DOI: 10.1111/j.1467-9868.2008.00663.x
发表时间: 2008-09-01
期刊: Journal of the Royal Statistical Society. Series B, Statistical methodology
影响因子: --
作者:
Banerjee S;Gelfand AE;Finley AO;Sang H
通讯作者: Sang H
DOI: 10.1214/15-aoas862
发表时间: 2015-12-01
影响因子: 1.8
作者:
Bradley, Jonathan R.;Holan, Scott H.;Wikle, Christopher K.
通讯作者: Wikle, Christopher K.
DOI: 10.1214/aos/1176350043
发表时间: 1986-09-01
影响因子: 4.5
作者:
BARRY, D
通讯作者: BARRY, D
DOI: 10.1016/0378-3758(78)90017-4
发表时间: 1978-01-01
影响因子: 0.9
作者:
BASU, D
通讯作者: BASU, D