Gaussian Process Planning with Lipschitz Continuous Reward Functions: Towards Unifying Bayesian Optimization, Active Learning, and Beyond

Gaussian Process Planning with Lipschitz Continuous Reward Functions: Towards Unifying Bayesian Optimization, Active Learning, and Beyond
复制标题

DOI:
10.1609/aaai.v30i1.10210
复制
发表时间:
2015-11
期刊:
--
影响因子:
--
通讯作者:
Chun Kai Ling;K. H. Low;Patrick Jaillet
Chun Kai Ling;K. H. Low;Patrick Jaillet
中科院分区:
其他
文献类型:
--
作者:
Chun Kai Ling;K. H. Low;Patrick Jaillet

文献摘要

被引文献

相似文献

本文提出了一种新的非近视自适应高斯过程规划(GPP)框架赋予了一个一般类的Lipschitz连续奖励函数,可以统一一些主动学习/传感和贝叶斯优化标准,并提供从业者一些灵活性,以指定他们所需的选择定义新的任务/问题。特别是,它利用了一个原则贝叶斯顺序决策问题框架,共同和自然优化的勘探开发权衡。一般来说,由于候选观测的不可数集合,所得到的诱导GPP策略不能精确地导出。因此,我们在这里的工作的一个关键贡献在于利用Lipschitz连续性的奖励函数来解决非近视自适应ε-最优GPP(ε-GPP)的政策。为了在真实的时间内进行规划,我们进一步提出了一个渐进最优的,分支定界的性能保证的任意时刻的变体的ε-GPP。我们实证证明了我们的ε-GPP政策的有效性,其随时变化的贝叶斯优化和能量收集任务。
This paper presents a novel nonmyopic adaptive Gaussian process planning (GPP) framework endowed with a general class of Lipschitz continuous reward functions that can unify some active learning/sensing and Bayesian optimization criteria and offer practitioners some flexibility to specify their desired choices for defining new tasks/problems. In particular, it utilizes a principled Bayesian sequential decision problem framework for jointly and naturally optimizing the exploration-exploitation trade-off. In general, the resulting induced GPP policy cannot be derived exactly due to an uncountable set of candidate observations. A key contribution of our work here thus lies in exploiting the Lipschitz continuity of the reward functions to solve for a nonmyopic adaptive epsilon-optimal GPP (epsilon-GPP) policy. To plan in real time, we further propose an asymptotically optimal, branch-and-bound anytime variant of epsilon-GPP with performance guarantee. We empirically demonstrate the effectiveness of our epsilon-GPP policy and its anytime variant in Bayesian optimization and an energy harvesting task.