Optimal Sequential Exploration: Bandits, Clairvoyants, and Wildcats

Optimal Sequential Exploration: Bandits, Clairvoyants, and Wildcats
复制标题

最优顺序探索:强盗、千里眼和野猫

DOI:
10.1287/opre.2013.1164
复制
发表时间:
2013
期刊:
Oper. Res.
影响因子:
--
通讯作者:
James E. Smith
James E. Smith
中科院分区:
--
文献类型:
--
作者:
David B. Brown;James E. Smith

文献摘要

被引文献

相似文献

本文的动机是制定北海油气田勘探的最佳政策。我们应该先在哪里钻孔?接下来我们在哪里钻?在这个问题和许多其他问题中,我们面临着收入(例如,立即在具有最大期望值的地点钻探)和学习(例如,在提供有价值信息的地点钻探)之间的权衡,这可能会在未来带来更大的收入。这些“顺序勘探问题”类似于多臂老虎机问题,但概率依赖性起着关键作用:钻探地点的结果揭示了有关邻近目标的信息。良好的勘探政策将利用这些信息的披露。我们为顺序探索问题开发启发式策略,并用最优策略性能的上限来补充这些启发式策略。我们首先将目标分组为可管理大小的集群。启发式方法源自将这些集群视为独立的模型。上限是通过假设每个簇具有有关所有其他簇的结果的完美信息而给出的。该分析在很大程度上依赖于强盗超级过程的结果,这是多臂强盗问题的概括。我们使用蒙特卡罗模拟评估启发式和界限,在北海示例中,我们发现启发式策略几乎是最优的。
This paper was motivated by the problem of developing an optimal policy for exploring an oil and gas field in the North Sea. Where should we drill first? Where do we drill next? In this and many other problems, we face a trade-off between earning (e.g., drilling immediately at the sites with maximal expected values) and learning (e.g., drilling at sites that provide valuable information) that may lead to greater earnings in the future. These “sequential exploration problems” resemble a multiarmed bandit problem, but probabilistic dependence plays a key role: outcomes at drilled sites reveal information about neighboring targets. Good exploration policies will take advantage of this information as it is revealed. We develop heuristic policies for sequential exploration problems and complement these heuristics with upper bounds on the performance of an optimal policy. We begin by grouping the targets into clusters of manageable size. The heuristics are derived from a model that treats these clusters as independent. The upper bounds are given by assuming each cluster has perfect information about the results from all other clusters. The analysis relies heavily on results for bandit superprocesses, a generalization of the multiarmed bandit problem. We evaluate the heuristics and bounds using Monte Carlo simulation and, in the North Sea example, we find that the heuristic policies are nearly optimal.