Multi-Armed Bandits for Human-Machine Decision Making
Multi-Armed Bandits for Human-Machine Decision Making
复制标题
DOI:
10.1109/icassp.2018.8461843
复制
发表时间:
2018-04
期刊:
影响因子:
--
通讯作者:
Paul B. Reverdy;Vaibhav Srivastava
中科院分区:
文献类型:
--
作者:
Paul B. Reverdy;Vaibhav Srivastava
Building an integrated human-machine decision-making system requires developing effective interfaces between the human and the machine. We develop such an interface by studying the multi-armed bandit problem, a simple sequential decision-making paradigm that can model a variety of tasks. We construct Bayesian algorithms for the multi-armed bandit problem, prove conditions under which these algorithms achieve good performance, and empirically show that, with appropriate priors, these algorithms effectively model human choice behavior; the priors then form a principled interface from human to machine. We take a signal processing perspective on the prior estimation problem and develop methods to estimate the priors given human choice data.