Multi-Armed Bandits for Human-Machine Decision Making

Multi-Armed Bandits for Human-Machine Decision Making
复制标题

DOI:
10.1109/icassp.2018.8461843
复制
发表时间:
2018-04
期刊:
2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Paul B. Reverdy;Vaibhav Srivastava
Paul B. Reverdy;Vaibhav Srivastava
中科院分区:
其他
文献类型:
--
作者:
Paul B. Reverdy;Vaibhav Srivastava

文献摘要

被引文献

相似文献

建立一个集成的人机决策系统需要在人和机器之间开发有效的接口。我们通过研究多臂强盗问题,一个简单的顺序决策模式,可以模拟各种任务,开发这样的接口。我们构造贝叶斯算法的多臂强盗问题,证明条件下,这些算法实现良好的性能,并实证表明,与适当的先验,这些算法有效地模拟人类的选择行为,先验,然后形成一个原则性的接口,从人到机器。我们采取信号处理的角度来看先验估计问题,并开发方法来估计先验的人类选择数据。
Building an integrated human-machine decision-making system requires developing effective interfaces between the human and the machine. We develop such an interface by studying the multi-armed bandit problem, a simple sequential decision-making paradigm that can model a variety of tasks. We construct Bayesian algorithms for the multi-armed bandit problem, prove conditions under which these algorithms achieve good performance, and empirically show that, with appropriate priors, these algorithms effectively model human choice behavior; the priors then form a principled interface from human to machine. We take a signal processing perspective on the prior estimation problem and develop methods to estimate the priors given human choice data.