Online Batch Decision-Making with High-Dimensional Covariates

Online Batch Decision-Making with High-Dimensional Covariates
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
ChiHua Wang;Guang Cheng
ChiHua Wang;Guang Cheng
中科院分区:
其他
文献类型:
--
作者:
ChiHua Wang;Guang Cheng

文献摘要

被引文献

相似文献

我们提出并研究了一类新的算法,顺序决策,互动与\textit{一批用户}同时,而不是\textit{一个用户}在每个决策时期。这种类型的批量模型的动机是互动营销和临床试验,其中一组人同时接受治疗,并在下一阶段的决策之前收集整个组的结果。在这种情况下,我们的目标是根据观察到的高维用户协变量分配一批治疗,以最大限度地提高治疗效果。我们提供了一个解决方案,命名为\textit{Teamwork LASSO Bandit algorithm},通过在整个决策过程中在团队阶段和自私阶段之间切换,解决了批处理版本的探索-利用困境。这是可能的基础上的统计特性LASSO估计的治疗效果,适应一系列的批次观察。在一般情况下,最优分配条件的速率,提出描绘的勘探和开发权衡的数据收集方案,这是足够的LASSO,以确定观察到的用户协变量的最佳处理。给出了该算法的期望累积遗憾的上界。
We propose and investigate a class of new algorithms for sequential decision making that interacts with \textit{a batch of users} simultaneously instead of \textit{a user} at each decision epoch. This type of batch models is motivated by interactive marketing and clinical trial, where a group of people are treated simultaneously and the outcomes of the whole group are collected before the next stage of decision. In such a scenario, our goal is to allocate a batch of treatments to maximize treatment efficacy based on observed high-dimensional user covariates. We deliver a solution, named \textit{Teamwork LASSO Bandit algorithm}, that resolves a batch version of explore-exploit dilemma via switching between teamwork stage and selfish stage during the whole decision process. This is made possible based on statistical properties of LASSO estimate of treatment efficacy that adapts to a sequence of batch observations. In general, a rate of optimal allocation condition is proposed to delineate the exploration and exploitation trade-off on the data collection scheme, which is sufficient for LASSO to identify the optimal treatment for observed user covariates. An upper bound on expected cumulative regret of the proposed algorithm is provided.