Incentivizing Bandit Exploration: Recommendations as Instruments
Incentivizing Bandit Exploration: Recommendations as Instruments
复制标题
激励 Bandit 探索:建议作为工具
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Zhiwei Steven Wu
中科院分区:
文献类型:
--
作者:
Daniel Ngo;Logan Stapleton;Nicole Immorlica;Vasilis Syrgkanis;Zhiwei Steven Wu
We study a multi-armed bandit learning setting where a social planner incentivizes a set of heterogeneous agents to efficiently explore the set of available arms. At each round, an agent arrives with their unobserved private type that determines both their prior preferences across the actions as well as their action-independent confounding shift in the rewards. The planner provides the agent with an arm recommendation that may alter their belief and incentivize them to explore potentially sub-optimal arms. Under this setting, we provide a novel recommendation mechanism that views the planner’s recommendations as a form of instrumental variables (IV) that only affect agents’ arm selection but not the observed rewards. We construct such IVs by carefully mapping the history–the interactions between the planner and the previous agents–to a random arm recommendation. Despite the unobserved confounding shift in the rewards, the resulting IV regression provides reliable estimates on the mean rewards of the actions and enables the social learning process to minimize regret over the long term.