Incentivizing Bandit Exploration: Recommendations as Instruments

Incentivizing Bandit Exploration: Recommendations as Instruments
复制标题

激励 Bandit 探索:建议作为工具

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Zhiwei Steven Wu
Zhiwei Steven Wu
中科院分区:
--
文献类型:
--
作者:
Daniel Ngo;Logan Stapleton;Nicole Immorlica;Vasilis Syrgkanis;Zhiwei Steven Wu

文献摘要

被引文献

相似文献

我们研究了一个多臂老虎机学习环境,其中社会规划者激励一组异质代理有效地探索可用的手臂集。在每一轮中,代理都会带着他们未被观察到的私人类型到达,这决定了他们在行动中的先验偏好以及他们在奖励中与行动无关的混杂转变。规划者向代理提供手臂建议,这可能会改变他们的信念并激励他们探索潜在的次优手臂。在此设置下,我们提供了一种新颖的推荐机制,将规划者的推荐视为工具变量(IV)的一种形式,仅影响代理的手臂选择,但不影响观察到的奖励。我们通过仔细地将历史(规划者和先前的代理之间的交互)映射到随机臂推荐来构建这样的 IV。尽管奖励中存在未观察到的混杂变化,但由此产生的 IV 回归提供了对行为平均奖励的可靠估计,并使社会学习过程能够最大限度地减少长期的遗憾。
We study a multi-armed bandit learning setting where a social planner incentivizes a set of heterogeneous agents to efficiently explore the set of available arms. At each round, an agent arrives with their unobserved private type that determines both their prior preferences across the actions as well as their action-independent confounding shift in the rewards. The planner provides the agent with an arm recommendation that may alter their belief and incentivize them to explore potentially sub-optimal arms. Under this setting, we provide a novel recommendation mechanism that views the planner’s recommendations as a form of instrumental variables (IV) that only affect agents’ arm selection but not the observed rewards. We construct such IVs by carefully mapping the history–the interactions between the planner and the previous agents–to a random arm recommendation. Despite the unobserved confounding shift in the rewards, the resulting IV regression provides reliable estimates on the mean rewards of the actions and enables the social learning process to minimize regret over the long term.