Estimation Considerations in Contextual Bandits

Estimation Considerations in Contextual Bandits
复制标题

上下文强盗中的估计注意事项

DOI:
--
复制
发表时间:
2017
期刊:
arXiv.org
影响因子:
--
通讯作者:
G. Imbens
G. Imbens
中科院分区:
--
文献类型:
--
作者:
Maria Dimakopoulou;S. Athey;G. Imbens

文献摘要

参考文献

被引文献

相似文献

上下文强盗算法寻求学习个性化的治疗分配策略,平衡探索和剥削。虽然已经提出了许多算法,但应用研究人员在各种方法中进行选择的指导意见很少。在计量经济学和统计学关于因果效应估计的文献的启发下,我们研究了对探索与开发框架的一个新的考虑,即目前的探索方式可能会导致在随后的学习阶段潜在结果模型估计的偏差和差异。我们利用参数和非参数统计估计方法以及因果效应估计方法来提出新的背景下的强盗设计。通过各种模拟,我们展示了替代设计选择如何影响学习绩效,并提供了为什么我们观察到这些影响的见解。
Contextual bandit algorithms seek to learn a personalized treatment assignment policy, balancing exploration against exploitation. Although a number of algorithms have been proposed, there is little guidance available for applied researchers to select among various approaches. Motivated by the econometrics and statistics literatures on causal effects estimation, we study a new consideration to the exploration vs. exploitation framework, which is that the way exploration is conducted in the present may contribute to the bias and variance in the potential outcome model estimation in subsequent stages of learning. We leverage parametric and non-parametric statistical estimation methods and causal effect estimation methods in order to propose new contextual bandit designs. Through a variety of simulations, we show how alternative design choices impact the learning performance and provide insights on why we observe these effects.
DOI: --
发表时间: 2018
期刊: Proceedings of the 21st International Conference on Artificial Intelligence and Statistics (AISTATS
影响因子: --
作者:
Kallus, Nathan;Zhou, Angela
通讯作者: Zhou, Angela
平衡的政策评估和学习
DOI: --
发表时间: 2018
期刊: Advances in neural information processing systems
影响因子: --
作者:
Kallus, Nathan
通讯作者: Kallus, Nathan