Stochastic Gaussian Process Model Averaging for High-Dimensional Inputs

Stochastic Gaussian Process Model Averaging for High-Dimensional Inputs
复制标题

DOI:
10.1109/wsc48552.2020.9384114
复制
发表时间:
2020-12
期刊:
2020 Winter Simulation Conference (WSC)
影响因子:
--
通讯作者:
Maxime Xuereb;S. Ng;Giulia Pedrielli
Maxime Xuereb;S. Ng;Giulia Pedrielli
中科院分区:
其他
文献类型:
--
作者:
Maxime Xuereb;S. Ng;Giulia Pedrielli

文献摘要

相似文献

许多统计学习方法在应用于大型高维数据集时表现出效率和准确性的损失。噪声数据加剧了这种损失。在本文中,我们专注于高斯过程(GP),这是一系列用于机器学习和贝叶斯优化的非参数方法。事实上,GP显示出难以随着输入数据大小和维度进行缩放。本文首次提出了随机GP模型平均(SGPMA)算法,以解决这两个挑战。SGPMA使用贝叶斯方法对多个预测因子进行加权,每个预测因子都使用初始数据集的独立子集进行训练(解决大数据集问题),并在原始空间的低维嵌入中定义(解决高维问题)。我们使用不同的输入大小和维度进行了几个实验。结果表明,我们的方法是上级朴素平均和嵌入的选择是至关重要的管理计算成本/预测精度的权衡。
Many statistical learning methodologies exhibit loss of efficiency and accuracy when applied to large, high-dimensional data-sets. Such loss is exacerbated by noisy data. In this paper, we focus on Gaussian Processes (GPs), a family of non-parametric approaches used in machine learning and Bayesian Optimization. In fact, GPs show difficulty scaling with the input data size and dimensionality. This paper presents, for the first time, the Stochastic GP Model Averaging (SGPMA) algorithm, to tackle both challenges. SGPMA uses a Bayesian approach to weight several predictors, each trained with an independent subset of the initial data-set (solving the large data-sets issue), and defined in a low-dimensional embedding of the original space (solving the high dimensionality). We conduct several experiments with different input size and dimensionality. The results show that our methodology is superior to naive averaging and that the embedding choice is critical to manage the computational cost / prediction accuracy trade-off.