An Actor-Critic Contextual Bandit Algorithm for Personalized Interventions using Mobile Devices

An Actor-Critic Contextual Bandit Algorithm for Personalized Interventions using Mobile Devices
复制标题

使用移动设备进行个性化干预的演员批评家上下文强盗算法

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
S. Murphy
S. Murphy
中科院分区:
--
文献类型:
--
作者:
Ambuj Tewari;S. Murphy

文献摘要

参考文献

被引文献

相似文献

适应性干预(AI)根据用户的持续表现和不断变化的需求,个性化干预的类型、模式和剂量。即时自适应干预(JITAI)利用现代移动设备提供的实时数据收集和通信能力,实时适应和提供干预措施。尽管JITAI越来越受到临床和行为科学家的欢迎,但在构建高质量的数据库JITAI方面缺乏方法学指导仍然是推进JITAI研究的一个障碍。在本文中,我们首次尝试通过将实时定制干预措施的任务表述为上下文强盗问题来弥合这种方法上的差距。然而,对可解释性的关注导致我们以不同于现有的网络应用程序(如广告或新闻文章放置)的表述方式来表述这个问题。我们将奖励函数(“批评家”)参数化与随机策略(“参与者”)的低维参数化分开选择。我们提供了一个在线演员评论算法来指导JITAI的构建和改进。给出了行动者评价算法的渐近性质,包括奖励参数和JITAI参数的一致性和收敛速度,并通过数值实验进行了验证。据我们所知,这是演员-评论家架构第一次应用于背景土匪问题。
An Adaptive Intervention (AI) personalizes the type, mode and dose of intervention based on users’ ongoing performances and changing needs. A Just-In-Time Adaptive Intervention (JITAI) employs the real-time data collection and communication capabilities that modern mobile devices provide to adapt and deliver interventions in real-time. The lack of methodological guidance in constructing databased high quality JITAI remains a hurdle in advancing JITAI research despite the increasing popularity JITAIs receive from clinical and behavioral scientists. In this article, we make a first attempt to bridge this methodological gap by formulating the task of tailoring interventions in real-time as a contextual bandit problem. However, interpretability concerns lead us to formulate the problem differently from existing formulations intended for web applications such as ad or news article placement. We choose the reward function (the “critic”) parameterization separately from a lower dimensional parameterization of stochastic policies (the “actor”). We provide an online actor-critic algorithm that guides the construction and refinement of a JITAI. Asymptotic properties of actor-critic algorithm, including consistency and rate of convergence of reward and JITAI parameters are provided and verified by a numerical experiment. To the best of our knowledge, our is the first application of the actor-critic architecture to contextual bandit problems.
DOI: 10.1007/s13142-011-0021-7
发表时间: 2011-03
影响因子: 3.6
作者:
Riley, William T.;Rivera, Daniel E.;Atienza, Audie A.;Nilsen, Wendy;Allison, Susannah M.;Mermelstein, Robin
通讯作者: Mermelstein, Robin