Nonparametric Bayesian Aggregation for Massive Data

Nonparametric Bayesian Aggregation for Massive Data
复制标题

DOI:
--
复制
发表时间:
2015-08
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Zuofeng Shang;Botao Hao;Guang Cheng
Zuofeng Shang;Botao Hao;Guang Cheng
中科院分区:
其他
文献类型:
--
作者:
Zuofeng Shang;Botao Hao;Guang Cheng

文献摘要

相似文献

我们为一般类别的非参数回归模型开发了一组可扩展的贝叶斯推理程序。具体来说,对从海量数据集中随机划分的每个子集分别进行非参数贝叶斯推理,然后将获得的局部结果聚合成全局对应结果。该聚合步骤是明确的,不涉及任何额外的计算成本。通过仔细的划分,我们表明我们的聚合推理结果获得了预言规则,因为它们相当于直接从整个数据中获得的结果(这在计算上是令人望而却步的)。例如,聚合的可信球在拥有与预言球相同的半径的同时,实现了理想的可信度水平和频率覆盖率。
We develop a set of scalable Bayesian inference procedures for a general class of nonparametric regression models. Specifically, nonparametric Bayesian inferences are separately performed on each subset randomly split from a massive dataset, and then the obtained local results are aggregated into global counterparts. This aggregation step is explicit without involving any additional computation cost. By a careful partition, we show that our aggregated inference results obtain an oracle rule in the sense that they are equivalent to those obtained directly from the entire data (which are computationally prohibitive). For example, an aggregated credible ball achieves desirable credibility level and also frequentist coverage while possessing the same radius as the oracle ball.