Bandit Learning with Delayed Impact of Actions

Bandit Learning with Delayed Impact of Actions
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Wei Tang;Chien-Ju Ho;Yang Liu
Wei Tang;Chien-Ju Ho;Yang Liu
中科院分区:
其他
文献类型:
--
作者:
Wei Tang;Chien-Ju Ho;Yang Liu

文献摘要

被引文献

相似文献

我们考虑具有延迟影响的随机多臂老虎机 (MAB) 问题。在我们的设置中,过去采取的行动会影响未来的手臂奖励。这种行动的延迟影响在现实世界中很普遍。例如,某个社会群体的人偿还贷款的能力可能取决于历史上该群体批准贷款申请的频率。如果银行继续拒绝弱势群体的贷款申请,可能会形成反馈循环,并进一步损害该群体获得贷款的机会。在本文中,我们在多臂强盗的背景下阐述了行动的这种延迟和长期影响。我们概括了老虎机设置,以编码由于学习过程中的动作历史而产生的这种“偏差”的依赖性。目标是随着时间的推移最大化收集的效用,同时考虑到历史行为的延迟影响所产生的动态。我们提出了一种算法,可以实现 $\tilde{\mathcal{O}}(KT^{2/3})$ 的遗憾,并显示匹配的遗憾下限 $\Omega(KT^{2/3})$,其中 $K$ 是臂数,$T$ 是学习范围。我们的结果通过添加处理具有长期影响的行为的技术来补充强盗文献,并对设计公平算法产生影响。
We consider a stochastic multi-armed bandit (MAB) problem with delayed impact of actions. In our setting, actions taken in the past impact the arm rewards in the subsequent future. This delayed impact of actions is prevalent in the real world. For example, the capability to pay back a loan for people in a certain social group might depend on historically how frequently that group has been approved loan applications. If banks keep rejecting loan applications to people in a disadvantaged group, it could create a feedback loop and further damage the chance of getting loans for people in that group. In this paper, we formulate this delayed and long-term impact of actions within the context of multi-armed bandits. We generalize the bandit setting to encode the dependency of this"bias"due to the action history during learning. The goal is to maximize the collected utilities over time while taking into account the dynamics created by the delayed impacts of historical actions. We propose an algorithm that achieves a regret of $\tilde{\mathcal{O}}(KT^{2/3})$ and show a matching regret lower bound of $\Omega(KT^{2/3})$, where $K$ is the number of arms and $T$ is the learning horizon. Our results complement the bandit literature by adding techniques to deal with actions with long-term impacts and have implications in designing fair algorithms.