Human-inspired algorithms for search A framework for human-machine multi-armed bandit problems

Human-inspired algorithms for search A framework for human-machine multi-armed bandit problems
复制标题

受人类启发的搜索算法 人机多臂老虎机问题的框架

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
Paul B. Reverdy
Paul B. Reverdy
中科院分区:
--
文献类型:
--
作者:
Paul B. Reverdy

文献摘要

参考文献

被引文献

相似文献

搜索是一种无处不在的人类活动。这是对我们在日常生活中寻求完成的任务中固有的不确定性的理性回应,从检索信息到做出重要决定。工程师们已经开发了许多工具来执行自动搜索,但许多任务都有太多的不确定性,这些工具在没有人工干预的情况下无法充分执行。因此,此类任务的工程解决方案由人机混合系统组成,在该系统中,人类主管与自动化工具进行交互,并做出高层决策来指导它们。在这种情况下,需要新的严格的人类决策模型来促进人机系统的原则性设计。在本文中,我们建立了一个严格的搜索任务中人类决策行为的模型。我们使用机器学习文献中的多臂强盗问题来形式化地建模搜索,这允许我们推导出最优决策性能的界限。我们将重点放在空间搜索上,为此我们引入了空间多臂强盗问题。我们从神经科学和机器学习的文献中扩展了启发式方法,建立了几个人类决策行为模型,并证明了其中一个模型(UCL)达到最优性能的条件。我们研究了空间多臂匪徒问题中的人类受试者数据,并表明人类在这个问题上的表现可以分为几类。一些人类在多臂匪徒问题上的表现优于标准算法,我们将其归因于人类对空间搜索具有良好的直觉。我们证明了UCL模型可以通过调整模型参数来获得落入不同类别的性能。模型参数将人类的直觉量化,并使其可用于人机系统。通过将UCL模型与统计学文献中的广义线性模型联系起来,我们给出了UCL模型的参数估计。UCL模型与估值器一起代表了人类决策的对象-观测器对,可用于系统设计。最后,我们考虑了一个所谓的“满意”目标,作为标准多臂强盗问题的最大化目标的替代。我们根据这一新目标推导出性能界限,并开发出一种实现最优性能的算法。
Search is a ubiquitous human activity. It is a rational response to the uncertainty inherent in the tasks we seek to accomplish in our daily lives, from retrieving information to making important decisions. Engineers have developed numerous tools to perform automated search, but many tasks have too much uncertainty for these tools to perform adequately without human intervention. Engineering solutions to such tasks therefore consist of human-machine hybrid systems, where human supervisors interact with automated tools and make high-level decisions to guide them. Novel rigorous models of human decision making in such situations are required to facilitate the principled design of human-machine systems. In this thesis, we develop a rigorous model of human decision-making behavior in search tasks. We formally model search using the multi-armed bandit problem from the machine learning literature, which allows us to derive bounds on optimal decision-making performance. We focus on spatial search, for which we introduce the spatial multi-armed bandit problem. We develop several models of human decision-making behavior in this problem by extending heuristics from the neuroscience and machine learning literatures, and prove conditions under which one model (UCL) achieves optimal performance. We study human-subject data from a spatial multi-armed bandit problem and show that human performance in this problem falls into several categories. Some humans outperformed standard algorithms for multi-armed bandit problems, which we attribute to humans having good intuition for spatial search. We show that the UCL model can achieve performance that falls in the different categories by tuning the model parameters. The model parameters quantify a human’s intuition and make it available to a humanmachine system. We develop a parameter estimator for the UCL model by relating it to the Generalized Linear Model from the statistics literature. The UCL model together with the estimator represent a plant–observer pair for human decision making which can be used for system design. Finally, we consider a so-called “satisficing” objective as an alternative to the maximizing objective of the standard multi-armed bandit problem. We derive performance bounds in terms of this new objective and develop an algorithm that achieves optimal performance.
DOI: 10.1371/journal.pcbi.1004237
发表时间: 2015-06
影响因子: 4.3
作者:
Wilson RC;Niv Y
通讯作者: Niv Y