Online Decision Making with High-Dimensional Covariates

Online Decision Making with High-Dimensional Covariates
复制标题

DOI:
10.1287/opre.2019.1902
复制
发表时间:
2020-01-01
影响因子:
2.7
通讯作者:
Bayati, Mohsen
Bayati, Mohsen
中科院分区:
管理学3区
文献类型:
--
作者:
Bastani, Hamsa;Bayati, Mohsen

文献摘要

被引文献

相似文献

大数据使决策者能够在各个领域(例如个性化医学和在线广告)的个人级别量身定制决策。这样做涉及学习决策奖励模型,这是根据特定的协变量有条件的。在许多实际情况下,这些协变量是高维度的。但是,通常只有一小部分观察到的特征可以预测决策的成功。我们将这个问题提出为具有高维协变量的K臂上下文匪徒,并根据LASSO估计器提出了一种新的有效的强盗算法。我们证明,在协变量d中,算法的累积预期遗憾量表最多是多数级的。据我们所知,这是上下文强盗的第一个绑定。我们分析的关键步骤是证明了一种新的尾巴不平等,尽管非i.i.d,但仍保证了套索估计器的融合。土匪政策引起的数据。此外,我们通过在药物给药问题的简化版本上对算法进行评估来说明算法的实际相关性。患者的最佳药物剂量取决于患者的遗传特征和病历;不正确的初始剂量可能会导致不良后果,例如中风或出血。我们表明,我们的算法的表现优于现有的匪徒方法和医生,以正确给大多数患者给药。
Big data have enabled decision makers to tailor decisions at the individual level in a variety of domains, such as personalized medicine and online advertising. Doing so involves learning a model of decision rewards conditional on individual-specific covariates. In many practical settings, these covariates are high dimensional; however, typically only a small subset of the observed features are predictive of a decision's success. We formulate this problem as a K-armed contextual bandit with high-dimensional covariates and present a new efficient bandit algorithm based on the LASSO estimator. We prove that our algorithm's cumulative expected regret scales at most polylogarithmically in the covariate dimension d; to the best of our knowledge, this is the first such bound for a contextual bandit. The key step in our analysis is proving a new tail inequality that guarantees the convergence of the LASSO estimator despite the non-i.i.d. data induced by the bandit policy. Furthermore, we illustrate the practical relevance of our algorithm by evaluating it on a simplified version of a medication dosing problem. A patient's optimal medication dosage depends on the patient's genetic profile and medical records; incorrect initial dosage may result in adverse consequences, such as stroke or bleeding. We show that our algorithm outperforms existing bandit methods and physicians in correctly dosing a majority of patients.